Ollaya – Ollama for open-source, Jev-style decision models
Ollaya is an open-source local runtime for fast Jev-style decision models, compatible with TypeSafe.
Ollaya is an open-source local runtime for Jev-style decision models, including Convai Innovations' Laya family, that returns calibrated answers in a single forward pass. It reports about 8-10 ms median latency for a five-question HTTP request on an RTX 4090 and serves TypeSafe-compatible /v1/systemone and /v1/models endpoints, so TypeSafe Python SDK 0.7.1 works unchanged. Desktop, CLI, and Docker builds cover macOS, Windows, and Linux under Apache-2.0, with NVIDIA GPU acceleration when CUDA 13 and driver R580 or newer are present.
- Decision models answer in one pass, about 10 ms on an RTX 4090.
- The server matches TypeSafe's API and works with Python SDK 0.7.1.
- Open weights include English, multilingual, and router Laya models.
- Apache-2.0 builds cover desktop, CLI, systemd, and Docker.
Full article667 words · extracted from ollaya.dev · click to collapse
Ask typed questions about any text or JSON and get calibrated answers in milliseconds. Private, open source, on your own hardware.
ollaya run laya --preset triage \"I was charged twice this month and want a refund."| Question | Answer | Probability |
|---|---|---|
| intent | refund | 1.00 |
| is_urgent | no | 0.87 |
| frustration | 1.59 / 3 clearly annoyed | 0.36 |
| refund_requested | yes | 0.88 |
| churn_risk | no | 0.89 |
Fast
Decisions in milliseconds.
A decision model answers in a single forward pass, with no token-by-token generation. On your own GPU, a five-question request to Laya takes about 10 ms, end to end through the HTTP API.
- laya:multilingual8.1 ms
- laya:en9.6 ms
- gliclass14.7 ms
- nli20.4 ms
- decider:0.8b155 ms
- decider:2b190 ms
- TypeSafe Jevhosted API236–276 ms
Ollaya: median of a five-question request through the HTTP API on an NVIDIA RTX 4090 (laya in fp16, the others in fp32). Jev: median request latency of the hosted API in third-party benchmarks (AbdelStark/jev-benchmarks, nibzard/decision-model-benchmark), which includes the network. Setups differ, so read it as an order-of-magnitude comparison.
Drop-in compatible
Speaks TypeSafe's API.
Ollaya serves /v1/systemone and /v1/models with TypeSafe's request and response shapes. The official TypeSafe Python SDK 0.7.1 works unchanged against a local server.
Request
# Point the TypeSafe SDK at Ollaya
export TYPESAFE_BASE_URL=http://localhost:11435
export TYPESAFE_API_KEY=local # any value works
export TYPESAFE_DEFAULT_MODEL=laya
# …or call the compatible endpoint directly
curl http://localhost:11435/v1/systemone -d '{
"model": "laya",
"state": "Can I get an invoice for last month?",
"questions": {
"intent": {
"type": "choice",
"instructions": "What does the customer want?",
"criteria": {
"invoice": "Needs an invoice or receipt",
"refund": "Wants money back",
"other": "Anything else"
}
}
}
}'Response
{
"model": "laya:en",
"answers": {
"intent": {
"type": "choice",
"choice": "invoice",
"confidence": 0.9547,
"probabilities": {
"invoice": 0.9698,
"refund": 0.0172,
"other": 0.013
}
}
},
"usage": {
"input_tokens": 43,
"output_tokens": 0
}
}Open models
Open weights, ready to pull.
Start with Laya from Convai Innovations: an English model, a 100+ language model, a model fine-tuned for typed decisions, and a router that picks for you.
Your data stays yours
Private by default.
Tickets, emails and user messages are often the most sensitive data you have. With Ollaya they are scored where they already live.
Platforms
Runs where you work.
A desktop app and a command line for macOS, Windows and Linux, and a Docker image for servers. Every model runs on the CPU; an NVIDIA GPU on Linux, in WSL 2 or in Docker takes a request down to milliseconds.
| Platform | Desktop app | Command line | GPU |
|---|---|---|---|
| macOSApple silicon | Desktop appMenu bar app.dmg | Command lineInstall script | GPUCPU only |
| Windows10 and 11, x64 | Desktop appDesktop app.exe or .msi | Command linePowerShell script | GPUCPU onlyNVIDIA via WSL 2 |
| Linuxx86-64 | Desktop appDesktop appAppImage, .deb, .rpm | Command lineInstall scriptsystemd service | GPUNVIDIA, CUDA 13 |
| LinuxARM64 | Desktop appNot available | Command lineInstall scriptsystemd service | GPUCPU only |
| WSL 2Linux on Windows | Desktop appNot available | Command lineInstall scriptSame as Linux | GPUNVIDIA, CUDA 13 |
| Dockeramd64 and arm64 | Desktop appNot available | Command lineImage on GHCR | GPUNVIDIA, CUDA 13:cuda image, amd64 |
NVIDIA GPUs need driver R580 or newer; the installers fetch the CUDA libraries only when they find one. On Apple, AMD and Intel GPUs, models run on the CPU.
Get up and running in minutes.
One binary, one command: ollaya run laya.
macOS, Windows, Linux and Docker · Apache-2.0 · GitHub