Maaz
All guides

/laya

Quick decisions over any text, without a giant chatbot.

A small open model that answers typed questions about text (which team? how urgent? yes or no?) in one pass. You run it yourself.

What it is

Laya is a decision engine. A chatbot writes an answer word by word; Laya scores a fixed set of answers in a single pass. You give it some text (an email, a support ticket, a JSON document) and typed questions:

  • choice: which option fits, like billing, technical or other;
  • score: where it sits on a scale, from “not urgent” to “blocking”;
  • yes/no: with the probability that the answer is yes.

It's an open alternative to Jev, and the README compares the two on public datasets. You host it yourself.

Why it matters

  • It's fast. One question takes about 33 ms on a T4 GPU, or 7.2 ms each when you batch them. There are no tokens to pay for.
  • There's nothing to parse. You get a label and a probability back, not free text, so it can't ramble or make things up.
  • It handles 100+ languages. A router sends non-English text to the multilingual model automatically.
  • You can fine-tune it on your own data. On the author's benchmark, fine-tuning raised accuracy from 0.362 to 0.766. Point Claude Code at the fine-tuning notebook in the repo and it can set up the run.

Get started

Python 3.10 or newer:

Terminal
python -m pip install laya
Python
from laya import Router

router = Router()  # downloads a model the first time
questions = {
    "department": {"type": "choice", "instructions": "Which department should handle this?",
                   "criteria": {"billing": "invoices, payments, refunds",
                                "technical": "bugs, outages, system errors",
                                "other": "everything else"}},
    "churn_risk": {"type": "noul", "instructions": "Does the user threaten to cancel?"},
}
result = router.predict("We were billed twice for March. Refund it or we cancel.", questions)
print(result["answers"]["department"]["choice"])  # billing

Want to see it first? Try the demo on Hugging Face.

Before you use it

  • 33 ms is on a GPU. On a laptop CPU it's slower, so measure on your own hardware.
  • Out of the box it's a starting point. The base model scored 0.362 on the author's benchmark. Plan to fine-tune it before you trust it with real routing.
  • Long documents need a setting. The multilingual model stops reading at 1,024 tokens unless you pass max_len=8192, and accuracy drops past about 4,000 tokens.
  • On an Intel Mac, a plain install won't run. Use the pinned versions in the README's install notes.

Get the next guide by email.

One email when a new guide or launch goes up. No spam, and you can unsubscribe from any email.