EMA Lightning is an 8.6M-parameter Apache-2.0 Turkish TTS model reporting 0.92% WER on Freya-TR-Eval.
EMA Lightning is a new Apache 2.0 Turkish text-to-speech release: a 5.6M-parameter diffusion-transformer acoustic model plus a 3M-parameter vocoder, 8.6M parameters and about 34 MB total. On Freya-TR-Eval it reports 0.92% word error rate, lower than ElevenLabs v4, Gemini 3.8, and the 2.38B-parameter Trendyol-TTS. On an RTX 4090 it claims first audio in 3.86 ms and 440× real-time speed, or 1,316× with batching, with streaming, offline CPU or GPU inference, and roughly $0.0085 per million characters.
Total size is 8.6M parameters, about 34 MB, under Apache 2.0.
Freya-TR-Eval word error rate is 0.92%, below ElevenLabs v4, Gemini 3.8, and Trendyol-TTS.
First audio in 3.86 ms and 440× real time on an RTX 4090; 1,316× with batching.
Runs fully offline on GPU or CPU and can stream while text is still generating.
Full article2,823 words · extracted from huggingface.co · click to collapse

# ⚡ EMA Lightning
**Tiny, fast and accurate Turkish text to speech.** A 5.6M-parameter DiT acoustic model and a 3M-parameter vocoder: 8.6M in total, about 34 MB. Apache 2.0.
- **Most accurate on Freya-TR-Eval:** 0.92% WER, the lowest of every system measured, including ElevenLabs v4, Gemini 3.8 and the 2.38B-parameter Trendyol-TTS.
- **Fast:** first audio in 3.86 ms and 440× faster than real time on an RTX 4090; 1,316× with batching.
- **Streams:** audio starts while the rest of the text is still being generated.
- **Cheap:** about $0.0085 per million characters on a rented RTX 4090.
- **Runs anywhere:** GPU or CPU, fully offline. No API key, and no text or audio leaves the machine.
- **One model serves many callers:** call it from as many threads as you like; its scheduler, Playhead, runs everyone's work together on the GPU.
**Figure 1.** Text is encoded per letter, placed on a word–letter timeline by a duration predictor, turned into one condition per frame by a windowed aligner, generated as latents in four steps by a DiT, and decoded to 48 kHz audio. Details under [Architecture](#architecture).
## Listen
Every clip below is made with `say()` at seed 0, straight from the text shown.
**Model sadece 8.6 milyon parametre ve çok hızlı.**
**Sabahın erken saatlerinde liman henüz uyanmamıştı. Balıkçılar ağlarını sessizce topluyor, martılar ise teknelerin etrafında dönerek şanslarını deniyordu.**
Make one `EMA` and keep it for the whole program. There are two ways to get speech out of it:
| | `say()` | `stream()` |
|---|---|---|
| Gives you | the whole audio, once it's ready | the audio piece by piece, as it's made |
| Takes | one text, or a list of texts | one text per call |
| Use it for | files, voiceovers, batch jobs | playing as you go: speakers, phone calls, live apps |
### `say()`: a whole text, back when it's ready
```python
from ema_lightning import EMA
tts = EMA() # downloads ema.pt and decoder.pt, uses the GPU if there is one
speech = tts.say("Merhaba, size nasıl yardımcı olabilirim?", path="merhaba.wav")
print(speech.duration, speech.sample_rate) # seconds of audio, 48000
```
`say()` returns a `Speech`: the `audio` (a float32 NumPy array, mono, between -1 and 1), its `sample_rate`, its `duration` in seconds, and the `seed` that made it. With `path`, it also writes a WAV file. A list is made in shared batches, one `Speech` per text, in order; with `path`, the clips go into that folder as `0.wav`, `1.wav`, …:
```python
speeches = tts.say(["Günaydın.", "Siparişiniz yola çıktı.", "İyi günler dileriz."], path="clips")
```
### `stream()`: hear it while it's being made
```python
import sounddevice as sd # pip install sounddevice
with sd.OutputStream(samplerate=48000, channels=1, dtype="float32") as speaker:
for chunk in tts.stream("Merhaba! Bu ses siz dinlerken üretiliyor. İlk kelimeyi duyduğunuzda, "
"cümlenin geri kalanı çoktan hazır."):
speaker.write(chunk)
```
The first chunk is one second of audio and is ready in about 4 ms on a GPU. The rest follows four seconds at a time, far faster than it plays, so playback never waits. Leave the loop early and the rest of that text's work is dropped. Joined together, the chunks are the same audio `say()` gives.
`stream()` takes one text per call. To stream many texts at once, call it once per text, each from its own thread. All the calls share the model, and Playhead runs them together on the GPU:
```python
from concurrent.futures import ThreadPoolExecutor
texts = ["Merhaba, size nasıl yardımcı olabilirim?", "Siparişiniz yola çıktı.", "Randevunuz onaylandı, görüşmek üzere."]
def stream_one(text):
chunks = []
for chunk in tts.stream(text):
chunks.append(chunk) # in a real app: send it to this caller right away
return chunks
with ThreadPoolExecutor(max_workers=len(texts)) as pool:
results = list(pool.map(stream_one, texts)) # one list of chunks per text, in the same order
```
### `.lightning()`: the fast path on NVIDIA GPUs
```python
tts = EMA().lightning()
```
Call it once at startup. It measures the best batch size for your GPU (once, then cached), compiles the model, records CUDA graphs, checks them against the plain path and prints when it's ready, with its first-audio time. Without it, or on a CPU, everything works the same, just slower.
| Option | Values | Default |
|---|---|---|
| `speed` | 0.25 to 4 | 1.0 |
| `seed` | non-negative integer; the same seed gives the same audio | random (returned in `speech.seed`) |
| `path` | `say()` only: a `.wav` file for one text, a folder for a list | none |
Any text is accepted, and text never raises. Invalid settings raise `ValueError` before any work starts.
## One model, many callers

Every `say()` and `stream()` call goes through Playhead, a scheduler inside `EMA`. A GPU making one sentence at a time is mostly idle, so Playhead collects what every caller needs at that moment and runs it together. It keeps two first-come-first-served queues: sentences waiting for the model, and audio windows waiting for the decoder. Each turn, it takes up to one batch from the front of each, makes them in one pass each on the GPU, and hands every window straight back to the call that asked for it, in order. Nobody's sentences are overtaken, a caller who hangs up leaves the queues, and a batch that fails fails only its own callers.
On one RTX PRO 6000, 64 streams started at the same instant made 887 seconds of audio in 0.74 seconds, and every one matched its text said alone.
## Benchmarks
Freya-TR-Eval, all 495 sentences, speed 1.0:
| # | System | Params | Size vs EMA | WER % | CER % | RTF | × real time | x real time (batch 64) | GPU memory |
- **Our runs** (every row without †): all 495 sentences, 3 seeds or takes, raw text in, an 8 kHz band, Whisper large-v3 with beam 5, and Freya's text normalisation. ElevenLabs uses its default voice (George) with default settings. Gemini uses the voice Kore.
- **† marks numbers as reported on the Anka TTS model card.**
- **All RTFs are measured on an RTX 4090**, as total generation time divided by total audio length, three passes after a warm-up, middle pass reported. Cloud APIs have no RTF or parameter count.
- **"±" is the spread across the three seeds.** EMA Lightning's spread comes from its lowest, highest and average result (0.79 to 1.05%); its CER has no spread.
- **‡** "Antalia 1" by Sezgin Saygili, Emre Kaplaner, Oncel Ozgul and Fikri San Koktas (Patientdesk.ai): [huggingface.co/cloud0day3/antalia-1](https://huggingface.co/cloud0day3/antalia-1). Three of its takes failed inside its own code and count as silence.
### Speed and cost
All on an RTX 4090 with `.lightning()`:
| | |
|---|---|
| One request | 440× real time (RTF 0.0023) |
| First audio, typical | 3.86 ms, from raw text, normalisation included |
| Batch of 64 | 1,316× real time; plain PyTorch 952× |
| A 6 h 7 min audiobook, batched | about 17 seconds |
| Cost at $0.74/hour | about $0.0085 per million characters (24,091 characters a second) |
| CPU | about 6× real time (RTF 0.166, measured on a cloud-container CPU) |
## Architecture
**Figure 1 (top of this card).** **(a)** Overview. Normalized Turkish text is encoded per letter; a duration predictor and a word–letter timeline place every output frame at a word and a position within it; a windowed aligner builds one condition per frame; a latent generator produces 64-dimensional latents at 25 Hz in four steps; a decoder upsamples them ×1920 to 48 kHz audio. **(b)** Windowed Gaussian aligner. Each frame attends only to letters of its own word and the adjacent words, with a Gaussian bias around its position. **(c)** DiT block with shared timestep modulation. One linear layer maps the timestep embedding to all modulation parameters; each of the six blocks adds a learned offset.
| Part | File | Parameters |
|---|---|---|
| Acoustic model: text encoder, durations, windowed aligner and the DiT latent generator | `ema.pt` | 5.6M |