Generated by All in One SEO v4.9.9, this is an llms.txt file, used by LLMs to index the site. # Matthew L. Jockers "Everything . . . in nature's vast workshop from the extinction of some remote sun to the blossoming of one of the countless flowers which beautify our public parks is subject to a law of numeration as yet unascertained.” (Joyce, Ulysses, 1922) ## Sitemaps - [XML Sitemap](https://www.matthewjockers.net/sitemap.xml): Contains all public & indexable URLs for this website. ## Posts - [Blog](https://www.matthewjockers.net/blog-posts/) - [The Seductive Allure of Generative AI for Storytelling](https://www.matthewjockers.net/2025/12/31/the-seductive-allure-of-generative-ai-for-storytelling/) - Before I ever became a computational humanist, I was a James Joyce scholar. What drew me to Joyce was not difficulty for its own sake, but his extraordinary intimacy with language and his willingness to follow the word wherever it led, even when that path violated every convention of narrative efficiency or reader ease. A - [Rethinking Range in the Age of Generative AI ](https://www.matthewjockers.net/2025/09/15/rethinking-range-in-the-age-of-generative-ai/) - I recently reread David Epstein’s Range (2019), a book I first encountered a few years ago when it seemed every leadership forum was extolling the virtues of grit, 10,000 hours, and early specialization. Epstein pushed back, persuasively arguing that generalists, not specialists, are better equipped to solve complex problems, especially in domains where rules are unclear and - [Sophie Online Editing Environment](https://www.matthewjockers.net/2008/01/30/sophie-online-editing-environment/) - This is worth a look for those thinking about online collaboration: http://www.sophieproject.org/ "Sophie's raison d'être is to enable people to create robust, elegant rich-media, networked documents without recourse to programming." There is a useful demo movie at http://www.futureofthebook.org/sophie/files/Making_a_Sophie_Book.html - [POS Tagging XML with xGrid and the Stanford Log-linear Part-Of-Speech Tagger](https://www.matthewjockers.net/2008/05/29/pos-tagging-xml-with-xgrid-and-the-stanford-log-linear-part-of-speech-tagger/) - Recently (4/2008) I had reason to Part-Of-Speech tag a whole mess of novels, around 1200. I installed the Stanford Tagger and ran my first job of 250 novels on an old G4 under my desk. Everything worked fine, but the job took six days. After that experience, I figured out how to utilize xGrid for - [Chronicle of Higher Education Article](https://www.matthewjockers.net/2008/07/28/chronicle-of-higher-education-article/) - This week the Chronicle of Higher Education ran an article written by Jennifer Howard about "literary geospaces." The article featured some work I have done mapping Irish-American literature using Google Earth (and also profiled the work of Janelle Jenstad who has been mapping early modern London).Photo by Noah Berger The bit about my Google Earth/Irish-American - [Executing R in Php](https://www.matthewjockers.net/2008/11/11/executing-r-in-php/) - For their final project, the students in my Introduction to Digital Humanities seminar decided to analyze narrative style in Faulkner's Sound and the Fury. In addition to significant off-line analysis, we are building a web-based application that allows visitors to compare the different sections of the novel to each other and also to new, unseen - [Machine-Classifying Novels and Plays by Genre](https://www.matthewjockers.net/2009/02/13/machine-classifying-novels-and-plays-by-genre/) - In the post that follows here, I describe some recent experiments that I (and others) have conducted. The goal of these experiments was to accurately machine-classify novels and plays (Shakespeare's) by genre. One of the most interesting results ends up having more to do with feature extraction than classification algorithm Background Several weeks ago, Mike - [Is it the Joyce Industry or the Shakespeare Industry?](https://www.matthewjockers.net/2009/07/01/is-it-the-joyce-industry-or-the-shakespeare-industry/) - At the recent Digital Humanities Conference in Maryland, Matthew Wilkins and I got into a discussion about famous authors and the "industries" of scholarship that their works have inspired (see Matt's blog post about our discussion and his survey analysis of the MLA bibliography). The first time I ever heard the term "industry" used in - [65,000 Texts to Mine?](https://www.matthewjockers.net/2010/02/10/65000-texts-to-mine/) - A story in the Feb. 7th issue of the Telegraph reports that the British Library is going to make 65,000 first edition texts available for public download via Amazon's Kindle. This news is almost as exciting as Google's decision some years ago to partner with a consortium of big libraries in order to digitize all - [Analyze This (Page)](https://www.matthewjockers.net/2010/03/12/you-are-hereblogs-mjockersstanford-edus-blog-analyze-this-page-analyze-this-page/) - "TAToo" is a fun Flash widget developed by Peter Organisciak at the University of Alberta. Peter works under the supervision of Digital Humanists Par Excellence and TAPoR Gurus Geoffrey Rockwell and Stan Ruecker. The widget (just some embed-able code) does "layman's" text analysis on the web pages in which its code is embedded. I've added - [Who's Your DH Blog Mate: Match-Making the Day of DH Bloggers with Topic Modeling](https://www.matthewjockers.net/2010/03/19/whos-your-dh-blog-mate-match-making-the-day-of-dh-bloggers-with-topic-modeling/) - Social Networking for digital humanities nerds? Which DH bloggers are you most compatible with? Let's get the right nerds with the right nerds--match making made in digital humanities heaven. After seeing Stefan Sinclair's Voyeuristic analysis of the Day of DH Blog posts, I wrote and asked him how to get access to the "corpus" of - [Digital Humanities: Methodology and Questions](https://www.matthewjockers.net/2010/04/23/digital-humanities-methodology-and-questions/) - Students in our new Literature Lab doing what English Majors do! Folks keep expressing concern about the future of the humanities, and the "need" for a next big thing. In fact, the title of a blog entry in the April 23, 2010 New York Times takes it for granted that the humanities need "saving." The - [Stalker (R) and the journey of the Jockers iPhone](https://www.matthewjockers.net/2010/04/27/stalker-r-and-the-journey-of-the-jockers-iphone/) - Lot’s of hoopla in the last few days over the discovery that the iPhone keeps a database of locations it has traveled. Wasn’t long before someone in the R community figured out how to tap into this file and with a mere two lines of code you can visualize where your phone has been on - [What is a Literature Lab: Not Grunts and Dullards](https://www.matthewjockers.net/2010/05/29/what-is-a-literature-lab-not-grunts-and-dullards/) - Yesterday's Chronicle of Higher Education ran an article by Marc Parry about the work we are doing here in our new Literature Lab with "big data." It's awfully nice to be compared to Lewis and Clark exploring the frontiers of literary scholarship, but I think the article fails to give due credit to the exceptional - [Panning for Memes](https://www.matthewjockers.net/2010/05/30/panning-for-memes/) - Over in the English Department Literature Lab, we have been experimenting with Topic Modeling as a means of discovering latent themes (aka topics) in a corpus of 19th century novels. Topic Modeling is an unsupervised machine learning process that employs Latent Dirichlet allocation. "It posits that each document is a mixture of a small number - [Auto Converting Project Gutenberg Text to TEI](https://www.matthewjockers.net/2010/08/26/auto-converting-project-gutenberg-text-to-tei/) - Those who do corpus level computational text analysis are always hungry for more and more texts to analyze. Though we've become adept at locating texts from a wide range of sources (our own institutional repositories as well as a number of other places including Google Books, the Internet Archive, and Project Gutenberg), we still face - [On Collaboration](https://www.matthewjockers.net/2010/10/11/on-collaboration/) - I've been hearing a lot about "collaboration," especially in the digital humanities. Lisa Spiro at Rice University has written a very informative post about Collaborative Authorship in the Humanities as well as another post providing Examples of Collaborative Digital Humanities Projects. Both of these posts are worth reading, and Spiro offers some well-thought out and - [SEASR Grant](https://www.matthewjockers.net/2010/11/02/seasr-grant/) - This month a group of researchers at Stanford, University of Illinois, University of Maryland, and George Mason were awarded a $790,000 grant from the Mellon Foundation to advance the prior work of the SEASR project. I'll be serving as the overall Project Director and as one of the researchers in the Stanford component of the - [Unigrams, and bigrams, and trigrams, oh my](https://www.matthewjockers.net/2010/12/22/unigrams-and-bigrams-and-trigrams-oh-my/) - I've been watching the ngrams flurry online, in twitter, and on various email lists over the last couple of days. Though I think there is great stuff to be leaned from Google's ngram viewer, I'm advising colleagues to exercise restraint and caution. First, we still have a lot to learn about what can and cannot - [On Pamphleteering and Pamphlet One](https://www.matthewjockers.net/2011/02/03/on-pamphleteering-and-pamphlet-one/) - Several months ago, a group of us from the Stanford Literary Lab wrote and sent out for review the article that now appears in Pamphlet 1 of the Lab. The article, titled "Quantitative Formalism: an Experiment" was submitted, peer-reviewed, and approved for publication in a prestigious literary journal. There was, however, a catch. The editors - [Kansas Irish Reprint](https://www.matthewjockers.net/2011/03/10/kansas-irish-reprint/) - Rowfont Press of Wichita, Kansas has just published a newly illustrated edition of Charles Driscoll's memoir Kansas Irish (with my Critical Introduction). The book is available at Amazon. Kansas Irish and the two sequels that follow provide the most complete and authentic rendering of Irish life on the American prairie in the 19th Century. - [On Distant Reading and Macroanalysis](https://www.matthewjockers.net/2011/07/01/on-distant-reading-and-macroanalysis/) - Earlier this week Kathryn Schultz of the New York Times published a rather provocative, challenging, and in my opinion under-researched and over-sensationalized article about my colleague Franco Morreti's work theorizing a mode of literary analysis that he has termed "distant-reading." Others have already pointed out some of the errors Schultz made, and I'm fairly certain - [Aberrant Adjectives in 19th Century Novels](https://www.matthewjockers.net/2011/09/18/aberrant-adjectives-in-19th-century-novels/) - I created the visualization below using Many Eyes and a data set derived from part-of-speech tagged novels from 19th century Britain. Found here are the 100 most "aberrant adjectives." Aberrant here is determined by selecting those words that have the greatest amount of usage deviation (measured by relative frequency) over a 13 decade time period. - [The LDA Buffet is Now Open; or, Latent Dirichlet Allocation for English Majors](https://www.matthewjockers.net/2011/09/29/the-lda-buffet-is-now-open-or-latent-dirichlet-allocation-for-english-majors/) - For my forthcoming book, which includes a chapter on the uses of topic modeling in literary studies, I wrote the following vignette. It is my imperfect attempt at making the mathematical magic of LDA palatable to the average humanist. Imperfect, but hopefully more fun than plate notation. . . . . . imagine a quaint - [Macroanalysis](https://www.matthewjockers.net/2012/05/28/macroanalysis/) - In preparation for the publication of my book (Macroanalysis: Digital Methods and Literary History, UIUC Press, 2013), I've begun posting some graphs and other data to my (new) website. To get the ball rolling, I have created an interactive "theme viewer" where visitors will find a drop down menu of the 500 themes I harvested - [Amicus Brief Filed](https://www.matthewjockers.net/2012/07/09/amicus-brief-filed/) - In the last chapter of forthcoming my book, I write about the challenges of copyright law and how many a digital humanist is destined to become a 19th-centuryist if the law isn't reformed to specifically allow for and recognize the importance of "non-expressive" use of digitized content.* This week the Amicus Brief that I co-authored - [DH2012 and the 2013 Busa Award](https://www.matthewjockers.net/2012/07/20/dh2012-and-the-2013-busa-award/) - I could not make it to the DH conference in Hamburg this year (though I did manage to appear virtually). As chair of the Busa Award committee I had the pleasure of announcing that Willard McCarty had won the award. Willard will accept the award in 2013 when DH meets at the University of Nebraska. - [Computing and Visualizing the 19th-Century Literary Genome](https://www.matthewjockers.net/2012/07/20/computing-and-visualizing-the-19th-century-literary-genome/) - I was unable to attend the DH 2012 meeting in Hamburg, but I recorded my paper as a screen cast, and my ever faithful colleague Glen Worthey kindly delivered it on my behalf. The full presentation can be viewed here as a QuickTime movie. - [Some Advice for DH Newbies](https://www.matthewjockers.net/2013/01/03/advice-for-dh-newbies/) - In preparation for a panel session at DH Commons today, I was asked to consider the question: "What one step would you recommend a newcomer to DH take in order to join current conversations in the field?" and then speak for 3 - 4 minutes. Below is the 5 minute version of my answer. . - [Thoughts on a Literary Lab](https://www.matthewjockers.net/2013/01/04/defining-a-literary-lab/) - [For the “Theories and Practices of the Literary Lab” roundtable at MLA yesterday, panelists were asked to speak for 5 minutes about their vision of a literary lab. Here are my remarks from that session--#147] I take the descriptor “literary lab” literally, and to help explain my vision of a literary lab I want to - [Unfolding the Novel](https://www.matthewjockers.net/2013/02/20/unfolding-the-novel/) - I'm excited to announce a new research project dubbed "Unfolding the Novel" (which is a play on both "paper" and "protein" folding). In collaboration with colleagues from the Stanford Literary Lab and Arizona State University and in partnership with researchers of the Book Genome project of BookLamp.com we have begun work that traces stylistic and - [Pronouns in 19th Century Fiction](https://www.matthewjockers.net/2013/02/22/pronouns-in-19th-century-fiction/) - Some folks I follow on Twitter (@scott_bot, @benmschmidt, @rayncordell, @foxyfolklorist, and others) were engaged in a conversation this week about the frequency of gendered pronouns in a corpus of 233 fairy tales from @foxyfolklorist's dissertation. For a bit of literary contextualization, I tweeted a bar graph showing the frequency of 13 pronouns in a corpus - ["A Matter of Scale"](https://www.matthewjockers.net/2013/03/28/a-matter-of-scale/) - Back in November, Julia Flanders and I were invited to stage a debate on the matter of “scale” in digital humanities research for the "Boston Area Days of DH" conference keynote: Julia was to represent the micro scale and I the macro. Julia and I met up during the MLA conference in January and began - ["Secret" Recipe for Topic Modeling Themes](https://www.matthewjockers.net/2013/04/12/secret-recipe-for-topic-modeling-themes/) - The recently (yesterday) published issue of JDH is all about topic modeling. It's a great issue, and it got me thinking about some of the lessons I have learned over seven or eight years of modeling literary corpora. One of the important things I have learned is that the quality of the final model (which - [25 days until the 2013 DH Fun Run](https://www.matthewjockers.net/2013/06/23/dh-fun-run-2013/) - Below is the route/elevation for the July 18, 2013 Unofficial (as in run at your own risk this has nothing to do with the conference) DH 2013 Fun Run. The route begins and ends on the north side of the UNL Student Union (fountain area). From campus we will go a few blocks east to - [Obi Wan McCarty](https://www.matthewjockers.net/2013/07/19/obi-wan-mccarty/) - [Below is the text of my introduction of Willard McCarty, winner of the 2013 Busa Award.] As the chair of the awards committee that selected Prof. McCarty for this award it is my pleasure to offer a few words of introduction. I'm going to go out on a limb this afternoon and assume that you - [Text Analysis with R for Students of Literature](https://www.matthewjockers.net/2013/09/03/tawr/) - [Update (9/3/13 8:15 CST): Contributors list now active at the main Text Analysis with R for Students of Literature Resource Page] Below this post you will find a link where you can download a draft of Text Analysis with R for Students of Literature. The book is under review with Springer as part of a - [A Festivus Miracle: Some R Bingo code](https://www.matthewjockers.net/2013/12/09/a-festivus-miracle-some-r-bingo-code/) - A few weeks ago my daughter's class was gearing up to celebrate the Thanksgiving Holiday, and I was asked to help prepare some "holiday bingo cards" for the kid's party. Naturally, I wrote a program in R for the job! (I know, I know, Maslow's hammer) Since I learned a few R tricks for making - [Characterization in Literature and the Macroanalysis Lab](https://www.matthewjockers.net/2014/01/08/characterization-in-literature-and-the-macroanalysis-lab/) - I have just posted the syllabus for my spring macroanalysis class focusing on Characterization in Literature. The class is experimental in many senses of the word. We will be experimenting in the class and the class will be an experiment. If all goes according to plan, the only thing about this class that will be - [Experimenting with "gender" package in R](https://www.matthewjockers.net/2014/02/25/experimenting-with-gender-package-in-r/) - Yesterday afternoon, Lincoln Mullen and Cameron Blevins released a new R package that is designed to guess (infer) the gender of a name. In my class on literary characterization at the macroscale, students are working on a project that involves a computational study of character genders. . . needless to say, the 'gender' package couldn't - [Simple Point of View Detection](https://www.matthewjockers.net/2014/04/06/simple-point-of-view-detection/) - [Note 4/6/14 @ 2:24 CST: oops, had a small error in the code and corrected it: the second if statement should have been "< 1.5" which made me think of a still simpler way to code the function as edited.] [Note 4/6/14 @ 2:52 CST: After getting some feedback from Jonathan Goodwin about Ford's The - [Text Analysis with R . . . coming soon.](https://www.matthewjockers.net/2014/04/21/text-analysis-with-r-coming-soon/) - My new book, Text Analysis with R for Students of Literature is due from Springer sometime in May. I got the cover proofs this week (below). Looking good:-) - [So What?](https://www.matthewjockers.net/2014/05/07/so-what/) - Over the past few days, several people have written to ask what I thought about the article by Adam Kirsch in New Republic ("Technology Is Taking Over English Departments The false promise of the digital humanities.") In short, I think it lacks insight and new knowledge. But, of course, that is precisely the complaint that - [A Novel Method for Detecting Plot](https://www.matthewjockers.net/2014/06/05/a-novel-method-for-detecting-plot/) - While studying anthropology at the University of Chicago, Kurt Vonnegut proposed writing a master's thesis on the shape of narratives. He argued that “the fundamental idea is that stories have shapes which can be drawn on graph paper, and that the shape of a given society’s stories is at least as interesting as the shape - [Reading Macroanalysis: The Hard Way!](https://www.matthewjockers.net/2014/06/12/reading-macroanalysis-the-hard-way/) - This past November, Judge Denny Chin ruled to dismiss the Authors Guild's case against Google; the Guild vowed they would appeal the decision and two months ago their appeal was submitted. I'll leave it to my legal colleagues to discuss the merit (or lack) in the Guild's various arguments, but one thing I found curious - [NHC Summer Institutes in Digital Humanities](https://www.matthewjockers.net/2014/12/09/nhc/) - I'm pleased to announce that Willard McCarty and I are leading a two-year set of summer institutes in digital humanities at the National Humanities Center. Here is the official announcement: "The first of the National Humanities Center’s summer institutes in digital humanities, devoted to digital textual studies, will convene for two one-week sessions, first in - [Plot Arcs (Schmidt Style)](https://www.matthewjockers.net/2015/01/05/plot-arcs-schmidt-style/) - A few weeks ago Ben Schmidt posted a provocative blog entry titled "Typical TV episodes: visualizing topics in screen time." It's worth a careful read. . . Ben began by topic modeling the closed captioning data from a series of popular TV series and then visualizing the ten most common topics over the time span - [Revealing Sentiment and Plot Arcs with the Syuzhet Package](https://www.matthewjockers.net/2015/02/02/syuzhet/) - Introduction This post is a followup to A Novel Method for Detecting Plot posted June 15, 2014. For the past few years, I have been exploring the relationship between sentiment and plot shape in fiction. Earlier today I posted an R package titled “syuzhet” to github. The package is designed to extract sentiment and plot - [The Rest of the Story](https://www.matthewjockers.net/2015/02/25/the-rest-of-the-story/) - My blog on February 2, about the Syuzhet package I developed for R (now available on CRAN), generated some nice press that I was not expecting: Motherboard, then The Paris Review, and several R blogs (Revolutions, R-Bloggers, inside-R) all featured the work. The press was nice, but I was not at all prepared for the focus to be placed on the one piece - [Some thoughts on Annie's thoughts . . . about Syuzhet](https://www.matthewjockers.net/2015/03/04/some-thoughts-on-annies-thoughts-about-syuzhet/) - Annie Swafford has raised a couple of interesting points about how the syuzhet package works to estimate the emotional trajectory in a novel, a trajectory which I have suggested serves as a handy proxy for plot (in the spirit of Kurt Vonnegut). Annie expresses some concern about the level of precision the tool provides and - [Is that Your Syuzhet Ringing?](https://www.matthewjockers.net/2015/03/09/is-that-your-syuzhet-ringing/) - Over the weekend, Annie Swafford published another installment in her ongoing critique of Syuzhet, the R package that I released in early February. In her recent blog post, an interesting approach for testing the get_transformed_values function is proposed[1]. Previously Annie had noted how using the default values for the low-pass filter may result in too much information loss, to which I - [A Ringing Endorsement of Smoothing](https://www.matthewjockers.net/2015/03/24/ringing_endorsement/) - On March 7, Annie Swafford posted an interesting critique of the transformation method implemented in Syuzhet. Her basic argument is that setting the low-pass filter too low may result in misleading ringing artifacts.[1] This post takes up the issue of ringing artifacts more directly and explains how Annie's clever method of neutralizing values actually demonstrates just - [My Sentiments (Exactly?)](https://www.matthewjockers.net/2015/04/01/my-sentiments-exactly/) - While developing the Syuzhet package--a tool for tracking relative shifts in narrative sentiment--I spent a fair amount of time gut-checking whether the sentiment values returned by the machine methods were a good match for my own sense of the narrative sentiment. Between 70% and 80% of the time, they were what I considered to be good sentence level matches. . - [Requiem for a low pass filter](https://www.matthewjockers.net/2015/04/06/epilogue/) - Ben Schmidt's and Scott Enderle's recent entries into the syuzhet discussion have beaten the last of the low pass filter out of me. I'm not entirely ready to concede that Fourier is useless for the larger problem, but they have convinced me that a better solution than the low pass is possible and probably warranted. What that better solution is remains an - [Cumulative Sentiments](https://www.matthewjockers.net/2015/04/28/cumulative-sentiments/) - This morning Andrew N. Jackson posted an interesting alternative to the smoothing of sentiment trajectories. Instead of smoothing the trajectories with a moving average, lowess, or, dare I say it, low-pass filter, Andrew suggests cumulative summing as a "simple but potentially powerful way of re-plotting" the sentiment data. I spent a little time exploring and thinking about his approach this - [That Sentimental Feeling](https://www.matthewjockers.net/2015/12/20/that-sentimental-feeling/) - Eight months ago I began a series of blog posts about my experiments using sentiment analysis as a proxy for plot movement. At the time, I had done a fair bit of anecdotal analysis of how well the sentiments detected by a machine matched my own sense of the sentiments in a series of familiar novels. In - [More Syuzhet Validation](https://www.matthewjockers.net/2016/08/11/more-syuzhet-validation/) - Back in December I posted results from a human validation experiment in which machine extracted sentiment values were compared to human coded values. The results were encouraging. In the spring, we mined the human coded sentences to help create a new sentiment dictionary that would, in theory, be more sensitive to the sort of sentiment - [Resurrecting a Low Pass Filter (well, kind of)](https://www.matthewjockers.net/2017/01/12/resurrecting/) - On April 6th, 2015, I posted Requiem for a low pass filter acknowledging that the smoothing filter as I had implemented it in the beta version of Syuzhet was not performing satisfactorily. Ben Schmidt had demonstrated that the filter was artificially distorting the edges of the plots, and prior to Ben's post, Annie Swafford had - [Syuzhet 1.0.4 now on CRAN](https://www.matthewjockers.net/2017/12/16/syuzhet-1-0-4/) - On Friday I posted an updated version of Syuzhet (1.0.4) to CRAN. This version has been available over on GitHub for a while now. In version 1.0.4, support for sentiment detection in several languages was added by using the expanded NRC lexicon from Saif Mohammed. The lexicon includes sentiment values for 13,901 words in each - [Revisiting Chapter Nine of Macroanalysis](https://www.matthewjockers.net/2019/03/18/new-network-viz/) - Back when I was working on Macroanalysis, Gephi was a young and sometimes buggy application. So when it came to the network analysis in Chapter 9, I was limited in terms of the amount of data that could be visualized. For the network graphs, I reduced the number of edges from 5,660,695 down to 167,770 ## Pages - [](https://www.matthewjockers.net/) - After 25 years, I left academia in 2021 to join Apple where I was hired as a distinguished research scientist working in ML-driven personalization and recommendations. Today I serve as a Senior Research Manger, overseeing the AIML personalization science teams for Apple App Store, Books, Podcasts, and Video. My last academic post was at Washington - [Publications](https://www.matthewjockers.net/publications/) - Books Jockers, Matthew L and Rosamond Thalken. Text Analysis with R for Students of Literature (2nd ed.). Springer. 2020. Archer, Jodie and Jockers, Matthew L. The Bestseller Code: Anatomy of the Blockbuster Novel. St. Martins Press USA. Penguin Press, UK. [Translations available in Chinese, Czech, German, Italian, Japanese, Portuguese, Russian, Spanish, and Turkish]. 2016 Jockers, Matthew L. Text Analysis with R for Students - [Blog](https://www.matthewjockers.net/academic-blog/) - [500 Themes](https://www.matthewjockers.net/macroanalysisbook/macro-themes/) - In Macroanalysis: Digital Methods and Literary History (UIUC Press, 2013), I explain how I extracted 500 themes from a corpus of 19th-century novels using Latent Dirichlet Allocation. On this page you can select any one of the 500 themes to see a cloud visualization of the key words and then a series of plots showing - [Confusion Matrices](https://www.matthewjockers.net/macroanalysisbook/confusion-matrices/) - Below are a series of confusion matrices referenced in Macroanalysis: Digital Methods and Literary History. UIUC Press, 2013. GENDER Gender Precision Recall M 0.801699717 0.771117166 F 0.76203966 0.793510324 DECADE Decade Precision Recall 1840 0.283783784 0.368421 1830 0.314285714 0.511628 1790 0.603305785 0.730000 1780 0.916666667 0.161765 1860 0.705882353 0.507042 1850 0.58 0.341176 1810 0.534591195 0.607143 1800 0.390625 - [Books](https://www.matthewjockers.net/books/) - BOOKS Archer, Jodie and Jockers, Matthew L. The Bestseller Code: Anatomy of the Blockbuster Novel. St. Martins Press. 2016 / Penguin Press, UK, 2016. See: http://www.archerjockers.com/books/ Jockers, Matthew L. Text Analysis with R for Students of Literature. Springer, 2014. Supporting materials ----. - [Courses](https://www.matthewjockers.net/courses/) - UNIVERSITY OF NEBRASKA Digital Literary Studies Reading Popular Literature: Contemporary Bestsellers “Macroanalysis.” (undergraduate / graduate course) “Microanalysis.”(undergraduate / graduate course) “Modern Fiction.” “Introduction to Literature.” “English Capstone: Ulysses.” “National Literatures: Irish Literature.” STANFORD UNIVERSITY “Ad Hoc Seminar in Corpus Stylistics.” 2009 (undergraduate / graduate course) “Literary Studies and the Digital Library.” 2009, 2010 (undergraduate / - [Workshop Materials](https://www.matthewjockers.net/materials/) - April 5, 2017: Vanderbilt University. Introduction to Text Analysis in R February 20-21, 2014: Michigan State University. Introduction to Text Analysis and Topic Modeling with R. September 12, 2013: University of Kansas. Introduction to Text Analysis and Topic Modeling with R. April 19, 2013: University of Milwaukee. Text Analysis and Topic Modeling in the Humanities. - [University of Chicago (2019)](https://www.matthewjockers.net/university-of-chicago-2019/) - Introduction to Text Analysis with R General Description: This workshop provides a basic introduction to text analysis using the R programming language. We will cover basic text processing, text ingestion, data preparation, and analysis. The main computing environment for the workshops will be R: “the open source programming language and software environment for statistical computing - [Text Analysis with R for Students of Literature](https://www.matthewjockers.net/text-analysis-with-r-for-students-of-literature/) - Text Analysis with R for Students of Literature provides a practical introduction to computational text analysis using the open source programming language R. Readers begin working with text right away and each chapter works through a new technique or process such that readers gain a broad exposure to core R procedures and a basic understanding - [The Bestseller Code](https://www.matthewjockers.net/books/the-bestseller-code/) - The Bestseller Code, Published September 20, 2016. St. Martins Press - [Noted](https://www.matthewjockers.net/noted/) - My Research in the News: 2017 Innovation of the Year Award 2016 "The Bestseller Code lays bare nothing less than the DNA of bestsellers, which makes Archer and Jockers the Watson and Crick of their age.” — The New Statesman "It's fun reading bestsellers after The Bestseller Code" — London Review of Books "may change - [University College London (May 10, 2017)](https://www.matthewjockers.net/university-college-london-may-10-2017/) - Introduction to Text Analysis with R General Description: This workshop provides a basic introduction to text analysis using the R programming language. We will cover basic text processing, data ingestion, data preparation, and analysis. The main computing environment for the workshops will be R: “the open source programming language and software environment for statistical computing - [Wesleyan (March 3, 2017)](https://www.matthewjockers.net/wesleyan-march-3-2017/) - Introduction to Text Analysis with R General Description: This workshop provides a basic introduction to text analysis using the R programming language. We will cover basic text processing, data ingestion, data preparation, and analysis. The main computing environment for the workshops will be R: “the open source programming language and software environment for statistical computing - [Vanderbilt (April 5, 2017)](https://www.matthewjockers.net/materials/vanderbilt-april-5-2017/) - Introduction to Text Analysis with R General Description: This workshop provides a basic introduction to text analysis using the R programming language. We will cover basic text processing, data ingestion, data preparation, and analysis. The main computing environment for the workshops will be R: “the open source programming language and software environment for statistical computing - [Year Two Instructions](https://www.matthewjockers.net/materials/national-humanities-center/year-two-instructions/) - Last year we got our hands dirty working through all of the chapters in Text Analysis with R for Students of Literature. This year, we'll do a "deep(ish) dive" into dplyr and xml2, two awesome R packages developed by R-Star, Hadley Wickham. In preparation for the year two R sessions, you should: Make sure that - [Year Two XML](https://www.matthewjockers.net/materials/national-humanities-center/year-two-xml/) - The Tragedy of Hamlet, Prince of Denmark Text placed in the public domain by Moby Lexical Tools, 1992. SGML markup by Jon Bosak, 1992-1994. XML version by Jon Bosak, 1996-1998. This work may be freely copied and distributed worldwide. Dramatis Personae CLAUDIUS, king of Denmark. - [National Humanities Center: June 8 - 12](https://www.matthewjockers.net/materials/national-humanities-center/) - Workshop Code Day One Day Two Day Three Day Four Day Five General Description: This workshop provides a practical introduction to text analysis using the R programming language. We will cover basic text processing, data ingestion, data preparation, and analysis. The main computing environment for the workshops will be R: “the open source programming language - [Day Five Code](https://www.matthewjockers.net/materials/national-humanities-center/day-five-code/) - source("code/day_five_functions.R") library(e1071) # Set the input Parameters input_dir - [Day Four Code](https://www.matthewjockers.net/materials/national-humanities-center/day-four-code/) - Functions ####################### # Day Four Functions # ####################### # Simple function for splitting a text by a delim split_text - [Day Three Code](https://www.matthewjockers.net/materials/national-humanities-center/day-three-code/) - # Day Three Code input_dir - [Day Two Code](https://www.matthewjockers.net/materials/national-humanities-center/day-two-code/) - # day_two_code.R ########################################### # There are two types of programmers: # # those who comment their code and those # # who are going to comment their code. # ########################################### ######################### # Day 2 Part 1 ######################### text_v - [Day One Code](https://www.matthewjockers.net/materials/national-humanities-center/day-one-code/) - ################################################ # Loading and Processing Moby Dick ################################################ text_v - [Harvard: April 3, 2015](https://www.matthewjockers.net/materials/harvard-april-3-2015/) - Introduction to Text Analysis with R General Description: This workshop provides a practical introduction to text analysis using the R programming language. We will cover basic text processing, data ingestion, data preparation, and analysis. The main computing environment for the workshops will be R: “the open source programming language and software environment for statistical computing - [Yale: December 5, 2014](https://www.matthewjockers.net/materials/yale-december-5-2014/) - Introduction to Text Analysis with R General Description: This workshop provides a practical introduction to text analysis using the R programming language. We will cover basic text processing, data ingestion, data preparation, and analysis. The main computing environment for the workshops will be R: “the open source programming language and software environment for statistical computing - [DH 2014: Introduction to Text Analysis and Topic Modeling with R](https://www.matthewjockers.net/materials/dh-2014-introduction-to-text-analysis-and-topic-modeling-with-r/) - Introduction to Text Analysis and Topic Modeling with R General Description: Introduction to Text Analysis and Topic Modeling with R is a two part workshops that will provide a practical introduction text analysis with a special emphasis on topic modeling. We will cover basic text processing, data ingestion, data preparation, and topic modeling. The main - [Macroanalysis](https://www.matthewjockers.net/macroanalysisbook/) - Ancillary Materials: Confusion Matrices Expanded Stop Words List The LDA Buffet: A Topic Modeling Fable 500 Labeled Themes with Graphs Color versions of Figures 9.3 and 9.4 Reviews and Comments Caldwell, Rob. Interview on the "207" at WCSH in Maine. April 18, 2014. Howard, Jennifer. Down in the Mines. The Times Literary Supplement. November 29, - [University of Gothenburg, Sweden: March 26, 2014](https://www.matthewjockers.net/materials/gothenburg/) - Introduction to Text Analysis and Topic Modeling with R General Description: Introduction to Text Analysis and Topic Modeling with R is a set of two workshops that will provide a practical introduction text analysis with a special emphasis on topic modeling. Taken together, the workshops will cover basic text processing, data ingestion, data preparation, and - [Michigan State University: February 20-21, 2014](https://www.matthewjockers.net/materials/msu/) - Introduction to Text Analysis and Topic Modeling with R General Description: Introduction to Text Analysis and Topic Modeling with R is a set of two workshops that will provide a practical introduction text analysis with a special emphasis on topic modeling. Taken together, the workshops will cover basic text processing, data ingestion, data preparation, and - [University of Kansas](https://www.matthewjockers.net/materials/university-of-kansas/) - Introduction to Text Analysis and Topic Modeling with R September 12, 2013 Instructor Contact: Matthew L. Jockers Email: mjockers@unl.edu Twitter: @mljockers Description: Introduction to Text Analysis and Topic Modeling with R will provide an introduction to computational text analysis and topic modeling in R. The course will cover basic text processing, data ingestion, data preparation, - [Workshop Code](https://www.matthewjockers.net/materials/uwm-2013/workshop-code/) - ############################################################### # Matthew L. Jockers # mjockers@unl.edu # Text Analysis and Topic Modeling in the Humanities Workshop # University of Wisconsin-Milwaukee # April 19, 2013 ############################################################### ############################################################### # SESSION ONE (9:00-10:15) ############################################################### ############################################# # 1.1 The R computing environment--what is R ############################################# ############################################# # 1.2 R console vs. RStudio show both quickly ############################################# ############################################# # - [UW-Milwaukee, 2013](https://www.matthewjockers.net/materials/uwm-2013/) - April 19, 2013: Text Analysis and Topic Modeling in the Humanities. University of Wisconsin-Milwaukee, 2013. Description: Text collections such as the HathiTrust Digital Library and Google Books have provided scholars in many fields with convenient access to their materials in digital form, but text analysis at the scale of millions or billions of words still - [2013 MLA/DH Commons](https://www.matthewjockers.net/materials/dh-commons-2013/) - January 3, 2013: Introduction to Topic Modeling Northeastern University DH Commons workshop. Workshop Materials (zip file) - [2013 DHWI ](https://www.matthewjockers.net/materials/dhwi-2013/) - January 7 - 11, 2013: Large Scale text Analysis with R. University of Maryland Digital Humanities Winter Institute. Syllabus Workshop Materials (zip file). The Day(s) in Code Day one code horde Day two code horde Day three code horde Day four code horde Functions File Some Reflections on the course content. . . @parezcoydigo Snuck - [DHWI: R Code Functions File](https://www.matthewjockers.net/materials/dhwi-2013/dhwi-r-code-functions-file/) - ############################################################### # mjockers unl edu # The Day in Code--DHWI Text Analysis with R. # Functions ############################################################### ####################################################################### # A Function to print a vector of file names in user friendly format ####################################################################### show.files - [DHWI: R Code Day Four](https://www.matthewjockers.net/materials/dhwi-2013/dhwi-r-code-day-4/) - ############################################################### # mjockers unl edu # The Day in Code--DHWI Text Analysis with R. # Day 4 ############################################################### # Today we pick up where we left of yesterday. . . # Review how mapply and xtabs work from the end of the day yesterday. ############################################################### # Here is the important code from yesterday again. . - [DHWI: R Code Day Three](https://www.matthewjockers.net/materials/dhwi-2013/dhwi-r-code-day-3/) - ############################################################### # mjockers unl edu # The Day in Code--DHWI Text Analysis with R. # Day 3 ############################################################### ############################################################ # Parsing XML ############################################################ setwd() library(XML) doc - [DHWI: R Code Day Two](https://www.matthewjockers.net/materials/dhwi-2013/dhwi-r-code-day-2/) - ############################################################### # mjockers unl edu # The Day in Code--DHWI Text Analysis with R. # Day 2 ############################################################### # Don't forget to set your working directory. . . . # Load Moby Dick File text - [DHWI: R Code Day One](https://www.matthewjockers.net/materials/dhwi-2013/dhwi-code/) - ############################################################### # mjockers unl edu # The Day in Code--DHWI Text Analysis with R. # Day 1 ############################################################### # Don't forget to set your working directory. . . . # Make moby.word.vector from Project Gutenberg Moby Dick text - [Expanded Stopwords List](https://www.matthewjockers.net/macroanalysisbook/expanded-stopwords-list/) - Below is the list of stop words I used in topic modeling a corpus of 3,346 works of 19th-century British, American, and Irish fiction. The list includes the usual high frequency words ("the," "of," "an," etc) but also several thousand personal names. 0, 1, 2, 3, 4, 5, 6, 7, 8, 9, a, aaron, abbey, - [The LDA Buffet: A Topic Modeling Fable](https://www.matthewjockers.net/macroanalysisbook/lda/) - . . . imagine a quaint town, somewhere in New England perhaps. The town is a writer’s retreat, a place they come in the summer months to seek inspiration. Melville is there, Hemingway, Joyce, and Jane Austen just fresh from across the pond. In this mythical town there is spot popular among the inhabitants; it - [Color Versions of Figures 9.3 and 9.4](https://www.matthewjockers.net/macroanalysisbook/color-versions-of-figures-9-3-and-9-4/) ## Categories - [Commentary](https://www.matthewjockers.net/category/commentary/) - [Text-Mining](https://www.matthewjockers.net/category/tm/) - [Tips and Code](https://www.matthewjockers.net/category/tandc/) - [DH 2013](https://www.matthewjockers.net/category/dh-2013/) - [Just for Fun](https://www.matthewjockers.net/category/just-for-fun/) - [R-Code](https://www.matthewjockers.net/category/r-code/) - [Academic](https://www.matthewjockers.net/category/academic/) - All legacy posts from my time as an academic.