Ollaya brings Ollama-style serving to small classification models
A new open-source runtime pulls and serves compact "decision models" locally, answering typed yes/no and scoring questions in milliseconds instead of round-tripping to a hosted LLM.
Ollaya is a new open-source CLI that pulls and runs small classification models locally, the same way Ollama does for chat models. Instead of generating text, it takes a piece of text or JSON plus a typed question and returns a calibrated answer: a choice, a score, or a yes/no. A single command installs it (curl -fsSL https://ollaya.dev/install.sh | sh) and it serves models over a local API compatible with the /v1/systemone wire format, so existing clients built against a hosted classifier can point at localhost by changing one environment variable.
The model lineup is deliberately small: laya (322m–421m parameters) for fast routing, decider (0.75b–4.2b) for higher accuracy, qwen3guard for safety screening, and a few others tuned for entailment and instruction-following classification. On an RTX 4090, laya answers in 8–10ms, against 236–276ms for a hosted equivalent. Everything is Apache-2.0.
Why it matters: if part of your pipeline calls a hosted LLM just to triage a ticket, route a request, or screen text for safety, that’s a classification problem, not a generation one. Running a purpose-built model on your own hardware removes a network round trip and a per-call bill for something a 400m-parameter model can do.
The caveat: commenters on the launch thread were quick to point out this isn’t a new technique, just text classification with a friendlier install path, and that the open models trail closed alternatives on harder queries. The value is in the packaging, not a research breakthrough.
ollaya run laya --preset triage "I was charged twice for my subscription this month and want a refund."