Tag Archives: corpora

The Lifespan of Words (three ways)

Getting ready for DH2017 this morning, I found myself curious about the lifespan of English words–when they come into the language and when they fall out. So I got all the earliest and latest attestation dates for all the words in OED3, and plotted them out. Here are three graphs (“visualizations,” if you like), all […]

Three conferences this summer

After a baby-related travelling hiatus of a couple three years, TLOW is hitting the road this summer, with stops at Ryerson University in Toronto (just barely down the road, really) at the end of May, for the Canadian Society for Digital Humanities meeting at CFHSS Congress; then off to Barbados and the University of the […]

One last round with metadata from Hathi and Underwood

In “Hathi’s Automatic Genre Classifier” and “Hathi Genre Again – Zero Recall“, I ran a couple of experiments comparing genre categories assigned by human taggers working on the Life of Words OED mark-up project to two sources of genre metadata associated with the HathiTrust Digital Library. The first post looked at data from the automatic […]

Shakespeare’s Earliest Citations in the OED

No author’s representation in the OED has received more comment than Shakespeare’s: if you ever come across a mention of OED citation evidence, more than likely it’s being used to substantiate (sometimes challenge or qualify) a claim that Shakespeare invented the most English words, or made up the most new meanings for existing words, or […]

OED Subject Matter

In my last post I described using HathiTrust’s Solr Proxy API to fetch Hathi genre metadata for OED quotations. But genre is not the only metadata that Hathi sends back down the intertubes when I ask it a question. For most works, I also get a Library of Congress Classification code for the volume. This […]

Hathi Genre Again – Zero Recall

In “Hathi’s Automatic Genre Classifier” [17.01.06] I compared the consolidated automatic genre metadata for a subset of HathiTrust Digital Library texts (available here) to the genre classifications arrived at for human-inspected works as part of the OED quotation tagging project under-way at The Life of Words. My process there was pretty closely supervised, but the […]

Hathi’s Automatic Genre Classifier

The HathiTrust Digital Library is a massive collection of digital books: As of 2017, it contains 5 billion pages from 15 million volumes (7 million titles). About 40% of these are public-domain works, meaning anyone can search and read them. Some of these have been marked for their textual genre. Here I do a little […]

Guest Post: Strong and Weak Genre Classification

Over the summer we’re featuring guest posts by Research Assistants at The Life of Words. Here Cosmin Dzsurdzsa – a 2nd year undergraduate in English at UW – thinks about moving from human intuition to computer rule-making in textual-genre classification: When trying to automate text classification algorithmically, one has to pay close attention to how […]

Vector Space and Poetic Logic

I’ve been spending the weekend experimenting with vector space modelling and poetic language. Vector space word embedding models use learning algorithms on very large corpora in order map a unique location in n-dimensional space to each token (=word) in the corpus. “N-dimensional space” is just a mathy-sounding way of saying that multiple (or n) features […]

Method as Tautology

Although it has been available for a while in the advanced access section of Digital Scholarship in the Humanities, and before that Literary and Linguistic Computing, my article on digital methods in literary research has recently been published in its final version. The full bibliographic details are: Williams, David-Antoine. “Method as Tautology in the Digital […]