animalese-js makes Animal Crossing style speech ("Animalese") in TypeScript. You give it text in Chinese, Japanese, Korean, or English. It gives you the babble that the villagers speak, in time with the dialogue box.
The work has two parts:
- An offline bake records the smallest sound units of each language with a TTS service, processes them, and packs them into voice banks.
- At runtime, the library splits the text into units, puts them on a timeline with a melody, and plays the banked units at a changed playback rate.
We measured the timing and the pitch values on recordings of Animal Crossing: New Horizons. The splitting rules are a good approximation, not the original algorithm of the game.
- Node.js 24 or later
- pnpm 11 or later
- An OpenAI-compatible TTS endpoint, only if you bake your own voice banks
The packages are not on npm yet. Use them from this repository.
pnpm install
pnpm devOpen the URL that Vite shows. The repository includes baked voice banks, so the playground works without a TTS endpoint.
import { createAnimalese } from 'animalese'
const animalese = await createAnimalese({ banks: '/banks/index.json' })
// Browsers start audio only after a user gesture.
button.onclick = () => animalese.say('哎呀,你来啦!', { voice: { preset: 'peppy' } })createAnimalese() returns three main functions:
say(text, options)plays the text and callsonEventas each character appears.plan(text, options)gives the tokens, the timeline, and the selected voice bank. It does not touch audio.render(text, options)gives a WAV file as aBlob.
Each call accepts a voice and a bank:
voiceis a preset, a preset with changes such as{ preset: 'cranky', speed: 7 }, or a fullVoiceOptionsobject.bankis the name of a recorded voice. If you do not give it, the library selects the voice that was recorded closest tovoice.baseHz.
The lower-level packages let you control each step:
import { schedule } from '@animalese/core'
import { analyze } from '@animalese/g2p'
import { BankLibrary, play } from '@animalese/web'
const tokens = analyze('今日はいい天気ですね')
const timeline = schedule(tokens, { preset: 'lazy' })
const library = await BankLibrary.fromIndex('/banks/index.json')
play(new AudioContext(), await library.load('nova', ['ja']), timeline)Node.js has no Web Audio. Use the renderer in @animalese/dsp instead:
import { readBank } from '@animalese/bake'
import { schedule } from '@animalese/core'
import { bankResolver, encodeWav, renderSchedule } from '@animalese/dsp'
import { analyze } from '@animalese/g2p'
const banks = [await readBank('apps/playground/public/banks/zh-nova')]
const timeline = schedule(analyze('你好呀'), { preset: 'peppy' })
const samples = renderSchedule(timeline, bankResolver(banks), 44100)
const wav = encodeWav({ samples, sampleRate: 44100 })The CLI does the same in one command:
pnpm say --bank apps/playground/public/banks/zh-onyx --text '嘿,过得还好吗?' --preset cranky --out hello.wavPut the TTS credentials in an env file. Then bake the languages and voices that you want:
pnpm bake --env-file .env --lang zh,ja,ko --voice nova,onyxThe baker reads the credentials in this order:
ANIMALESE_TTS_BASE_URL,ANIMALESE_TTS_API_KEY, andANIMALESE_TTS_MODELTESTING_AUDIO_TTS_*OPENAI_BASE_URLandOPENAI_API_KEY
The baker keeps each raw recording in .cache/tts. When you bake again with different settings, it sends requests only for units that are not in the cache. TTS services sometimes return silence for one syllable. The baker rejects these recordings and records them again.
This project uses the recipe of the Animalese tutorial by 重轻:
- Record the shortest sounds of a language, such as Chinese initials and finals.
- Cut the recording into single sounds.
- Trigger the sounds at random.
- Raise the pitch.
- Snap the pitch to a musical scale.
The tutorial does these steps by hand in a DAW. animalese-js does steps 1, 2, and the flat pitch offline. At runtime, it only splits text, plans the timeline and the melody, and changes the playback rate.
flowchart LR
subgraph Offline bake
A[Carrier text] --> B[TTS recording]
B --> C[Trim silence]
C --> D[Shorten consonant]
D --> E[Flatten pitch]
E --> F[Sprite and manifest]
end
subgraph Runtime
G[Text] --> H[G2P tokens]
H --> I[Text clock and voice clock]
I --> J[Melody]
J --> K[Playback]
end
F --> K
A voice bank holds every unit of one language, recorded with one TTS voice:
| Language | Units | Carrier text for the recording |
|---|---|---|
| Chinese | 402 toneless syllables | One common character for each syllable, from the first level of GB2312. The generator prefers characters with one reading, then the first tone. |
| Japanese | 101 kana morae, with yōon | The kana |
| Korean | 323 open syllables (19 initials × 17 vowels) | The syllable |
| English | No units of its own | English syllables use the nearest kana from the Japanese bank. |
The baker does these steps on each voiced unit:
- It trims the silence at the start and the end.
- It keeps a maximum of 30 ms of consonant before the voice starts.
- It makes the pitch flat with TD-PSOLA, at the reference pitch of the bank. TD-PSOLA keeps the length and the formants.
- It cuts the unit to a maximum length, adds fades, and sets the loudness.
- It packs all units into one
sprite.wav. Themanifest.jsonfile gives the position of each unit.
The reference pitch is the median pitch of the bank, rounded to a semitone. Because every unit has the same flat pitch, the runtime gets each note from one playback rate.
In the game, the dialogue box and the voice have different speeds. The text appears at the typing speed. The voice says a maximum number of units each second. On each tick, the voice says the latest character that appeared. The voice skips the characters that come faster than it can speak. It skips weak characters first, such as the Chinese neutral tone (的, 了, 吗) and English function words.
We measured these values on recordings of the Chinese and English versions of the game:
| Value | Chinese | English |
|---|---|---|
| Text speed | 11–13 characters/s | 45–70 letters/s |
| Voice speed | 8–10 units/s | 13–14 units/s |
| Units for each character or syllable | 0.6–0.8 | 0.75–1 |
Other measured values:
- One unit lasts 70–140 ms. Its pitch is almost flat, with a small fall at the end.
- The pitch changes much between villagers: about 114 Hz for the cranky 罗博 (Lobo), 330 Hz for Marshal, 445 Hz for Derwin, and 500 Hz for 小桃.
Because of these differences, the voice options give the pitch in Hz. A low voice uses a bank recorded with a low voice (onyx). The library does not play a high recording at a slower rate.
The recordings do not show if the game selects the skipped characters on purpose, or if the voice is only too slow. We did not measure Japanese and Korean yet, so they use the Chinese pacing.
| Package | Purpose |
|---|---|
animalese |
Entry point. createAnimalese() connects the analysis, the timeline, the bank selection, and playback. |
@animalese/core |
Types, voice presets, language pacing, the two clocks (revealTimes, voiceClock), the melody (shapeMelody), and schedule(). It has no dependencies. |
@animalese/g2p |
Language front ends: Chinese with pinyin-pro, Japanese with wanakana, Korean with es-hangul, and English syllables mapped to kana. It also routes mixed-script text and holds the unit inventories. |
@animalese/dsp |
Audio processing in plain TypeScript: WAV, resampling, YIN pitch tracking, TD-PSOLA, unit preparation, sprite packing, and offline rendering. |
@animalese/web |
BankLibrary (index, bank selection by pitch, cached loading), play(), render(), and renderWav() with Web Audio. |
@animalese/bake |
The baker: a cached recorder for OpenAI-compatible TTS, bakeBank(), readBank(), and the animalese-bake CLI. |
The playground uses Vite, React, React Router, Base UI, and animal-island-ui. It has three pages:
- 说话 (Speak): Type text and see the dialogue box appear with the voice. You can export WAV files.
- 词表 (Inventory): See each unit on the syllable chart of its language, and listen to it.
- 原理 (Pipeline): See each step of the pipeline, with controls. The offline steps run the real DSP code in the browser, on built-in recordings or on your microphone.
animalese-js is early, and its APIs can change. These limits are known:
- Japanese kanji do not have a dictionary yet. Each kanji maps to a fixed kana, so the rhythm is correct but the reading is not.
- If mixed text has hiragana, the library reads Han characters as Japanese. Otherwise, it reads them as Chinese. For text such as 東京タワー, give
language: 'ja'. - A TTS service can read a Chinese carrier character with a different pronunciation. We did not listen to all 402 carriers.
- The voice banks are uncompressed 16-bit WAV, about 9 MB for each voice.
This repository is a pnpm workspace:
pnpm install
pnpm test:run
pnpm typecheck
pnpm lint
pnpm buildpnpm lint runs moeru-lint, which runs oxlint and then ESLint. The ESLint configuration follows Project AIRI.
The table of contents in this README comes from doctoc. After you change the headings, run:
pnpm docs:tocResearch notes from the calibration (in Chinese) are in research/.
Note
This project is part of the Project AIRI ecosystem.
- 重轻's Animalese tutorial for the recipe
izure1/animalese-ttsfor the Japanese mora tableAcedio/animalese.jsandstefanlegg/animalese-webXinqwq/Animalese_Converter(ChineseGibberish)pinyin-pro,wanakana, andes-hangulanimal-island-uiand ChillRound