Navigation

Big (and Open) Data for Scholarship of All Sizes: A New Release of the HathiTrust Research Center Extracted Features Dataset

Questions or inquiries can be directed to HTRC Project Coordinator Ryan Dubnicek  (rdubnic2@illinois.edu)

December 5, 2016

HathiTrust today announces the release of a significantly expanded open dataset, the HathiTrust Research Center (HTRC) Extracted Features (EF) Dataset, Version 1.0. This dataset provides researchers with open access to data extracted from the full text of the HathiTrust Digital Library (HTDL) at an unprecedented scale. 

The Extracted Features Dataset opens the complete HathiTrust collection for investigations into historical and cultural trends, the rise and fall of topics within the corpus, and the evolution of words and writing structures in publications dating from the 16th to the late 20th century. The EF Dataset provides quantitative information about word and line counts, parts of speech, and other details within each page of every volume in the HTDL. In addition to these larger-scale investigations, the EF Dataset also allows researchers to closely analyze the contents of a given volume or subset of volumes.

The data is extracted from 13.7 million volumes found in the HTDL, representing over 5 billion pages consisting of over 2 trillion tokens (words). A preliminary release of the EF Dataset, drawn from a much smaller subset comprising only HathiTrust’s public domain collection, has already enabled novel research from scholars in economics, history, linguistics, literary studies and sociology, among other fields.

“The Extracted Features Dataset creates opportunities for scholarship and teaching that were previously impossible,” said J. Stephen Downie, co-director of HathiTrust Research Center and Associate Dean for Research and Professor at the School of Information Sciences, University of Illinois at Urbana-Champaign. “We look forward to seeing how the scholarly community takes advantage of the EF dataset in their research, labs, and classrooms.”

“We launched the HathiTrust Research Center to help researchers fully mine the entire collection of texts found in HathiTrust,” said Michael Furlough, HathiTrust’s executive director. “This release provides a novel and effective way to do so by generating relevant data from the entire corpus.”

Founded in 2008 and hosted at the University of Michigan, HathiTrust preserves and provides access to millions of digitized books and journals from the collections of more than 120 institutional academic and research partners via its certified trusted digital repository This searchable archive of published literature from around the world includes both in-copyright and public domain materials from mass digitization programs and partners’ local digitization initiatives.

The HathiTrust Research Center is an advanced research service of HathiTrust and a collaborative research center launched jointly by Indiana University and the University of Illinois.  The Research Center team strives to meet the technical challenges that researchers face when dealing with massive amounts of digital text, by developing cutting-edge software tools and cyberinfrastructure to enable advanced computational access to the growing digital record of human knowledge.

For more information about the Extracted Features Dataset and access to it, go to https://analytics.hathitrust.org/datasets. The HTRC EF Dataset is released under a Creative Commons CC-BY license. Download information can be found at the DOI in the formal dataset citation below:

Boris Capitanu; Ted Underwood; Peter Organisciak; Timothy Cole; M. Janina Sarol; J. Stephen Downie (2016): The HathiTrust Research Center Extracted Features Dataset. 1.0 [Dataset]. HathiTrust Research Center. Dataset. http://dx.doi.org/10.13012/J8X63JT3

 Questions? Please contact htrc-help@hathitrust.org.