Glossa's word lists are built separately for each of its five languages from open-licensed word lists, ranked by how often words are actually used, filtered through block lists, and corrected by human review. Each language gets two lists: a short grid list of common words that puzzles may require, and a much larger bonus list of real words the game accepts as extras. Nothing is translated between languages.
The name Glossa comes from the Greek word for "tongue" or "language", so we care about getting this part right. Here is how it works.
Why does a word game need two word lists?
A word puzzle has two jobs that pull in opposite directions.
- The answers must be fair. If a level requires a word, most players should know it. A required word nobody has heard of feels like a trick.
- The game must be generous. If you spell a real word and the game rejects it, that feels wrong too.
So every language in Glossa has:
- A grid list: the 8,000 most frequent words in English, Spanish, Dutch and French, and 9,000 in German. Only these can be required answers in a level.
- A bonus list: every real word of three to eight letters the source lists contain for that language, from roughly 66,000 words in Dutch to more than 160,000 in Spanish. Spell any of these and it counts as a bonus word.
- A block list: slurs and vulgar words, removed from both lists.
Words are three to eight letters long because the wheel never holds more than eight letters.
Where do the words come from?
Every source is open-licensed, pinned to an exact version and checked by checksum when it is downloaded, so the same inputs always produce the same lists.
| Use | Source | Licence |
|---|---|---|
| English, Spanish and French words | lorenbrichter/Words | CC0 1.0 |
| German words | enz/german-wordlist | CC0 1.0 |
| Dutch words | OpenTaal word list, © OpenTaal | CC BY 3.0 (also BSD-3-Clause) |
| Word frequency, all five languages | FrequencyWords, built from OpenSubtitles 2018 | CC BY-SA 4.0 |
| Block list starting point | LDNOOBW | CC BY 4.0 |
We picked word-game and spelling lists without share-alike terms on the word lists themselves. OpenTaal requires attribution, and the app credits it in Settings › About.
How is frequency used?
The frequency data counts how often words appear in film and TV subtitles, which is a decent stand-in for everyday spoken language. We use it only to rank and filter the words from the lists above. None of the frequency counts ship in the app.
How is the grid list chosen?
The grid list starts as the most frequent words in the bonus list, then passes four filters. These affect only which words can be required; a word that fails them can still be accepted as a bonus word.
- Nothing from the published block list. Some words on the block list turn out to be ordinary, such as the Dutch ASBAK (ashtray) or the French BAISER (to kiss). Review lets them back in as bonus words, but they are never required answers.
- No vowel-less or stuttered strings. Subtitles are full of HMM and AAAH. They are not puzzle answers.
- No English leakage. Subtitles in other languages contain English words. A word that is at least twenty times more frequent in English subtitles than in the language's own is left out of that language's grid list.
- No names or international tokens. A word used about as often in three other languages is probably a name or an international term. Everyday international words that clearly belong, such as TAXI and HOTEL, are kept by hand.
Reviewers also keep some words out of the grid that are real but not much fun as a required answer, such as words about murder or suicide. They remain valid bonus words; a puzzle just never asks for them.
How are accents and special letters handled?
Each language has a letter set and a set of rules for comparing words. Every word is converted to a standard "comparison form": upper case, with hyphens, spaces and apostrophes removed.
- English: A to Z. US and UK spellings are both accepted, so COLOR and COLOUR both count.
- Spanish: accents fold (ÁRBOL is compared as ARBOL), but Ñ stays its own letter, so PENA and PEÑA are different words.
- Dutch: A to Z, with accents folded. IJ is treated as two letters, I and J.
- German: Ä, Ö and Ü are distinct letters (BÄR is not BAR), and ß is written as SS, so STRASSE.
- French: accents fold, so ÉTÉ is compared as ETE. Hyphens are ignored.
The same rules run in the iPhone app, the Android app and the list builder, and shared test cases check that all three agree.
How does human review fit in?
Automatic filters get a list most of the way, not all of the way. Review lives in small override files, one per language, plus one for words every language should treat the same. They record decisions such as:
- allow: a real word the sources missed.
- block: an offensive word or word stem the starting list did not catch.
- block allow: a word on the published block list that is actually ordinary or clinical, accepted as a bonus word only.
- grid keep / grid deny: a word to keep in, or keep out of, the required answers.
The first pass was drafted by our team. Each language's grid list also goes to a native speaker for review, and their changes go into the same files. To be clear about scope: people review the grid lists and the override decisions, not every one of the hundreds of thousands of bonus words. The bonus lists are as good as their open sources plus our filters, and when a real word is missing, we add it.
Because the lists are rebuilt from the sources plus these files, every fix is a small, reviewable change, and rebuilding gives the same result every time.
How do the lists become levels?
Our level tools read the lists and build all 2,000 Journey levels per language, plus 100 levels per puzzle pack, ahead of time.
- Each level's wheel comes from the letters of a word on the grid list, with the most common words used first, so early levels feel familiar.
- The grid is filled with grid-list words the wheel can spell, favouring common ones and avoiding the same words over and over.
- Every other bonus-list word the wheel can spell is stored with the level as its bonus words.
- A separate checker rebuilds everything from the finished files and confirms each level is valid.
Because each level carries its own bonus words, the app does not need a dictionary on your phone. Word checks are instant and work offline. A dictionary fix reaches you with an app update.
Why are words not translated between languages?
A letter wheel is about letters, not meanings. Translate an English level into German and you get different letters, different lengths and a puzzle that no longer works. Common words also differ: what is everyday in Spanish is not simply the translation of what is everyday in English.
So each language has its own sources, its own frequency ranking, its own filters and its own levels. The only place a concept is shared is the Daily Emoji, where one emoji clue points to each language's own word, checked by a native speaker before release.
Want to put the lists to work? Try our tips for solving letter-wheel puzzles, or see the FAQ for more about the languages.
In short
- Each language has a grid list of common words (8,000, or 9,000 in German) that puzzles can require, and a far larger bonus list of real extras.
- Words come from open-licensed lists (CC0 lists, OpenTaal for Dutch), ranked by OpenSubtitles frequency data that never ships in the app.
- Block lists, frequency filters and per-language letter rules shape both lists; Ñ and German umlauts are distinct letters, other accents fold.
- People review the grid lists and override decisions, not every bonus word; fixes are small, reviewable changes.
- Levels are built ahead of time with their bonus words included, so the game works offline and nothing is translated between languages.