How it works

wordshift turns the frequency of words in published books into visual curves spanning more than two centuries. This page explains what you are looking at, how the tool works and how to get the most from it. No technical background needed.

The chart

When you search a word on wordshift, you see a curve running from 1800 to 2022. That curve shows how often the word appeared in published English books across those two centuries. A line that rises sharply means writers were suddenly reaching for that word far more often. A line that falls means it was fading from the conversation. A flat line near zero means it barely existed in print at all.

The shape of a word’s curve is its story.

The scale

Because some words appear millions of times in print and others barely at all, comparing them directly on the same chart would make most words invisible. wordshift solves this by indexing each word to its own peak. Whatever year a word appeared most often in the corpus is set to 100. Every other year is shown as a percentage of that peak.

This means you can put “pandemic” next to “cholera” on the same chart and see how their curves relate to each other, even though pandemic appears far more often in absolute terms. You are not comparing raw numbers. You are comparing the shape of each word’s history.

If you want to see the raw numbers, how often a word actually appeared per million words of text, switch to Actual frequency using the toggle above the chart.

The dots

Not every point on a curve is equally interesting. wordshift automatically detects the moments that stand out: the peaks, the troughs and the points where a word’s frequency changed significantly relative to its own history. These are marked with dots on the curve.

The detection works the same way geographers measure the significance of a mountain peak: by how much it rises above the surrounding terrain, not simply by how high it is. A modest peak in an otherwise flat landscape earns a dot. A high point that barely rises above a general plateau does not.

The moments

Click any dot and a moment opens below the chart. It shows the year, the word and a historical annotation explaining what was happening in English publishing at that point: the legislation, the crisis, the scientific discovery or the cultural shift that sent writers reaching for that word more often.

The annotations are generated by Claude, Anthropic’s AI language model, using the context of the specific word and year. They are honest about uncertainty. When a pattern looks like a data anomaly rather than a genuine historical event, the annotation says so.

Using wordshift

What to search

Any word or phrase in English. Single words work well. Two or three word phrases work if they were genuinely used as fixed expressions in print: “shell shock”, “climate change” or “the Turing test”. Very recent coinages, slang and internet-specific language often do not appear because the corpus covers published books, not social media or websites.

Searching two or three words at once reveals the most interesting stories: when one term displaced another, when two ideas rose and fell together or when a word peaked in one era and was replaced by another in the next. Try democracy and totalitarianism, consumption and tuberculosis or wireless and radio.

What the data actually is

wordshift draws on the Google Books Ngram corpus, a collection of over eight million digitised books in English covering more than two centuries of publishing. Think of it as a vast digital library, not of every book ever written but of a very large and carefully assembled sample of them.

The word corpus simply means a body of text used for analysis. The Google Books corpus is one of the largest ever assembled. Developed by researchers at Harvard and Google, it has been widely used by academics, journalists and language enthusiasts since its launch in 2010 and has been cited in peer-reviewed research across history, linguistics, psychology and the social sciences.

The books were scanned digitally, which means the text was read by a machine rather than typed by a human. This process is called optical character recognition, or OCR. It is highly accurate but not perfect. Occasionally a scanner misreads a letter or a character, which can make a word appear in a year it was never actually used. If you see a small spike for a word in an unexpected year, particularly before 1850 when scanning older materials is harder, that is usually a scanning error rather than genuine history. wordshift flags these where it can.

How accurate is it

wordshift uses the 2022 edition of the Google Books Ngram corpus, the most recent available. For most historical and cultural questions about English language publishing, the corpus is more than adequate and its patterns are consistent with what historians and linguists have documented through other means. That said, it is a sample, not a complete record. It over-represents certain publishers, genres and time periods and under-represents others. Books in languages other than English are not included. Self-published books and digital-only publications are largely absent.

This means wordshift shows you what was happening in the publishing conversation, in books, journals and printed literature, not in everyday spoken language or online writing. Treat the patterns as evidence worth investigating, not as definitive proof. The best use of wordshift is as a starting point for curiosity, not an ending point for argument.

Can I use this for research

Yes, with the caveats above in mind. The underlying data is widely cited in academic literature. The 2010 Science paper by Michel et al. that introduced the field of culturomics used the same corpus. If you cite wordshift in formal work, note the corpus version, the date range and the smoothing setting, all of which appear in the credit line below the chart. You can also download the data as a CSV file using the download button above the chart.

Frequently asked questions

Researchers and journalists use it to track how language has shifted around specific topics or events. Writers use it to check whether a word was in common use during a historical period they are writing about. Teachers use it to bring language history into the classroom. And a lot of people use it simply because the curves are surprising. wordshift is designed for anyone curious about language and history, regardless of background.

No. If you can read a line graph you can use wordshift. The indexed scale, the smoothing and the moment detection all work automatically. You search, the chart appears and you click the dots to read the story behind each inflection point.

Yes, with some caveats. The Google Books Ngram corpus has been widely used in academic research since its launch in 2010 and has been cited in peer-reviewed studies across history, linguistics, psychology and the social sciences. The 2010 Science paper by Michel et al. that introduced the field of culturomics used the same corpus and remains one of the most cited papers in digital humanities. If you cite wordshift in formal work, note the corpus version (2022), the date range and the smoothing setting used. All of these appear in the credit line below the chart.

Yes. Use the download button above the chart to export the data as a CSV file. The file includes a metadata header showing the words searched, the date range, the smoothing setting, the scale mode and the data source, so the file is self-describing if you open it later or share it.

This is almost always a scanning error. The books in the Google Books corpus were digitised using optical character recognition, a process where a machine reads printed text from a scanned page. OCR is highly accurate but not perfect, particularly with older typefaces and damaged pages. A misread character can produce a sequence of letters that matches a modern word, making it appear in years it was never actually used. Small isolated spikes in the early part of a curve, particularly before 1850, are usually this kind of artefact rather than genuine historical usage. wordshift flags these patterns in the moment annotations where it can identify them.

The moment annotations are generated by Claude, Anthropic’s AI language model. They draw on historical and publishing context for the specific word and year but are not verified against primary sources. Occasionally an annotation will misidentify the cause of an inflection point, particularly for words that have changed meaning over time or for patterns that are genuine corpus artefacts. If you spot an error, use the Flag issue button in the moments panel and we will review it.

Yes, up to three words or phrases at once. Use the plus button next to the search field to add a second or third word. Each word gets its own coloured curve and its own set of dots. Comparing two or three words often reveals the most interesting stories: when one term displaced another, when two ideas rose in tandem or when competing vocabularies crossed over. Try shell shock and PTSD or democracy and totalitarianism to see what this looks like in practice.

Use the share button above the chart to copy a link. The link encodes your search terms and settings so whoever opens it sees exactly the same chart. If you have a moment open, use the Share moment button in the moments panel to copy a link that opens directly on that specific dot and annotation.

Yes, wordshift is free to use with no registration required.

wordshift currently uses the 2022 edition of the Google Books Ngram corpus, the most recent version available. When Google releases a new corpus version we will evaluate updating the data. Any changes to the corpus version will be noted in the credit line below the chart.