
Who Laya is for#
Support teams routing and triaging incoming tickets
Laya's choice and score question types map directly to triage workflows: route to department, rate urgency, detect churn risk, flag refund requests. A single predict call answers all questions per ticket in under 40 ms. The multilingual checkpoint handles non-English tickets automatically, so a single model covers a global support queue without per-language routing logic.
Skip if:
Your support volume is low enough that keyword rules or a simple classifier already covers the routing. Laya's latency and accuracy advantages matter most at scale or with multilingual inputs where rules break down.
ML engineers building guardrail and moderation layers
The noul (yes/no) question type and calibrated probability output make Laya a direct guardrail component: flag requests for moderation, detect policy violations, and threshold on confidence. Because it runs in a single forward pass without generating text, it adds minimal latency to an existing inference pipeline as a pre- or post-processing step.
Skip if:
You need to detect nuanced semantic violations across documents consistently over 4,000 tokens. Accuracy on very long inputs becomes variable, and zero-shot performance on novel violation categories may require fine-tuning.
Teams migrating from TypeSafe Jev to self-hosted inference
laya-serve exposes the same POST /v1/systemone endpoint shape as TypeSafe Jev and accepts every Jev question type. If your team already uses the Jev SDK or direct API calls, switching to self-hosted Laya requires only a base URL change in your client configuration. Apache-2.0 licensing means no per-call fees and full control over the serving infrastructure.
Skip if:
You rely on Jev features outside the typed-decisions API shape, such as managed SLA guarantees or observability tooling built into the Jev platform. laya-serve covers the request/response contract but not managed operational services.
Data engineers classifying structured events in batch pipelines
Laya accepts JSON as the state input alongside plain text, so it can classify structured events, webhooks, and API payloads directly without extraction preprocessing. The batch CLI path (laya --batch FILE --predict) processes large files in one forward pass, measured at 2.6x the throughput of sequential predict calls on 20 tickets.
Skip if:
Your events carry structured fields that rule-based logic can already classify accurately. Laya's value is classifying unstructured or semi-structured text where deterministic rules are brittle or unmaintainable.
The problem it solves#
Building classification and triage logic on top of generative LLMs introduces compounding costs: the model runs slowly for pure classification tasks, produces free-form text you must parse back into structured types, and can hallucinate label values that look correct but are not. Lighter encoder-only classifiers avoid text generation but require a separate fine-tuned model for every question type, and they typically cover only the language they were trained on.
For teams handling incoming support tickets, moderation queues, or routing workflows across multiple languages, the operational challenge compounds further: maintaining separate models per language and question type, managing different latency profiles, and stitching together outputs from multiple inference calls instead of getting all answers from a single request.
How it solves it#
Single forward pass for all question types
Evaluates every question in the request against the input text in one encoder forward pass, returning choice, score, and yes/no answers in a single call. On a T4 GPU, one question resolves in 39.5 ms on the English checkpoint and 32.8 ms on the multilingual checkpoint. A batch of ten questions takes 72.3 ms on the multilingual model.
100+ language support with automatic routing
The built-in Router detects script and language in under 0.5 ms of pure Python and dispatches to the English checkpoint (ModernBERT-large) or the multilingual checkpoint (mmBERT-base) automatically. On the MASSIVE intent benchmark across 13 non-English languages, the Router achieves 0.451 accuracy by matching the multilingual checkpoint, well above the English-only model's 0.306 on the same set.
Calibrated probabilities via RLCD training
Trained with reinforcement learning against strictly proper scoring rules (RLCD), so reporting honest probabilities is the only way to maximize reward during training. The English checkpoint scores 0.860 on XNLI English; the multilingual checkpoint scores 0.731 on XNLI across 14 languages. Calibrated outputs let downstream logic threshold on confidence without additional calibration steps.
Fine-tuning for domain accuracy
The shipped checkpoints work zero-shot, but fine-tuning on domain-specific decisions produces substantial accuracy gains. On the typed-decisions benchmark of 2,000 decisions across four workflows, the fine-tuned laya-typed-decisions checkpoint scores 0.766 accuracy against 0.362 for the base English checkpoint zero-shot on the same set. A fine-tuning notebook runs the full loop on Kaggle's free 2x T4 GPUs.
Jev-compatible HTTP server
The laya-serve extra exposes the Router on the same POST /v1/systemone endpoint shape as TypeSafe Jev, accepting every Jev question type. Existing clients that point at Jev's API work against laya-serve by changing only the base URL. Configurable via LAYA_DEVICE and LAYA_PRELOAD environment variables for GPU deployment with preloaded checkpoints.
Framework and protocol integrations
Optional extras add first-class support for LangChain and LangGraph (laya[langchain]), LlamaIndex selectors (laya[llamaindex]), CrewAI routing (laya[crewai]), an MCP server (laya[mcp]), and ONNX Runtime for CPU-optimized deployment (laya[onnx]). Each integration plugs into existing orchestration pipelines without custom wrapper code.
Strengths and trade-offs#
Strengths
- No text generation means no parsing or hallucinationBecause Laya is an encoder-only model that returns typed answers directly, there is no generated text to parse and no risk of the model producing an invented label. This is structurally different from classifying with an LLM: the output is always one of the types you defined, at a calibrated probability. The model cannot generate a value outside the question schema.
- Apache-2.0 license covers all checkpoints and the server runtimeThe Apache-2.0 license applies to all three checkpoints and the laya-serve runtime. You can run Laya on your own infrastructure for any commercial use, modify the code, and redistribute it without licensing fees. No data leaves your environment unless you choose the Hugging Face Hub for initial checkpoint downloads.
- Router prevents silent accuracy collapse on non-Latin scriptsThe English checkpoint scores 0.000 accuracy on Khmer while reporting 0.952 confidence, meaning confidence gating cannot detect the failure. The Router identifies non-Latin scripts in under 0.5 ms and redirects to the multilingual checkpoint before the forward pass, eliminating this silent failure mode in multilingual pipelines without any caller-side configuration.
- Long-document support up to 8,192 tokens on the multilingual checkpointThe multilingual checkpoint extends to 8,192-token inputs with max_len=8192. At up to 4,000 tokens, accuracy holds at 16 to 18 of 20 correct on the benchmark. This covers most support tickets, emails, and documents without chunking or summarization as a preprocessing step.
Trade-offs
- -Accuracy becomes variable on documents over 4,000 tokensWhile laya-multilingual supports up to 8,192 tokens, the benchmark shows accuracy dropping from 16 to 18 of 20 correct in the 0 to 4,000 token range to 8 to 17 of 20 beyond that point. Long legal documents, technical reports, or verbose email threads near or above this threshold should be tested on domain data before deployment.
- -Cold checkpoint load adds seconds of latency on first requestLoading a checkpoint for the first time costs several seconds of setup: the median reload on CPU is 7.4 s; on a T4 GPU it is 10.3 s when max_loaded=1. A production server should preload checkpoints at startup with Router(preload=True) or LAYA_PRELOAD=1; without preloading, the first request to each language checkpoint absorbs this startup cost.
- -Zero-shot accuracy on specialized domain workflows is limitedThe shipped checkpoints work zero-shot, but the typed-decisions benchmark shows the base English checkpoint scoring only 0.362 accuracy on four specialized workflows, compared to 0.766 for the domain-fine-tuned version. Teams with industry-specific vocabulary or custom question types will likely need fine-tuning to reach production-grade accuracy on those workflows.
Laya vs alternatives#
Laya vs TypeSafe Jev
Both tools handle typed decisions over text using the same /v1/systemone API shape. The primary difference is deployment model: TypeSafe Jev is a managed commercial API service; Laya is Apache-2.0 licensed and runs on your own infrastructure.
| Feature | Laya | TypeSafe Jev |
|---|---|---|
| License | Apache-2.0 | Proprietary |
| Self-hosting | Yes | No |
| API endpoint | POST /v1/systemone | POST /v1/systemone |
| Languages | 100+ with auto-routing | Not publicly specified |
| Context window | Up to 8,192 tokens | Not publicly specified |
| Fine-tuning | Yes, notebooks provided | Not available |
| Pricing | Free on self-hosted infrastructure | Paid managed API |
Laya is the better choice when you need data residency guarantees, want to avoid per-call API fees at high volume, or need to fine-tune on proprietary domain data. Because laya-serve matches the Jev endpoint shape, existing integrations require only a base URL change to switch.
TypeSafe Jev is worth considering when you want a fully managed service with no infrastructure overhead. Running Laya on GPU hardware requires provisioning and maintaining a server; if your team lacks ML infrastructure experience or your usage volume makes managed API pricing affordable, Jev removes that operational work.
Migration from TypeSafe Jev
The laya-serve server accepts every TypeSafe Jev question type at the same POST /v1/systemone path. Teams using the Jev SDK or direct API calls can migrate by pointing their client at a self-hosted laya-serve instance. A TypeScript and Node.js SDK (laya-ts, published as npm install laya-ts) covers browser and server-side JavaScript integrations.
Quick start#
Install Laya with pip; add the serve extra to deploy a Jev-compatible HTTP server.
```bash
pip install laya
pip install "laya[serve]"
LAYA_DEVICE=cuda LAYA_PRELOAD=1 laya-serve
```What it's built on#
- Languages
- JavaScriptPythonTypeScript
- Frameworks
- FastAPIPyTorch
FAQ#
Is Laya free to use commercially on self-hosted infrastructure?
Yes. Laya is Apache-2.0 licensed, covering all three checkpoints and the laya-serve runtime. You can run it on your own infrastructure for any commercial purpose, including as an internal service or a public-facing API, without licensing fees. The Hugging Face Hub is used only for the initial checkpoint download; inference runs fully on your hardware after that.
How does Laya's accuracy compare to using an LLM for text classification?
On the MASSIVE intent benchmark for English, the Laya English checkpoint scores 0.783, comparable to results achievable with larger generative models, while running in 39.5 ms on a T4 GPU. On non-English languages the Router achieves 0.451 across 13 languages by dispatching to the multilingual checkpoint. Fine-tuning on domain data brings the typed-decisions checkpoint to 0.766 on the four standard workflows, compared to 0.362 for the base checkpoint zero-shot on the same set.
Can Laya replace TypeSafe Jev without changing client code?
For teams using the Jev HTTP API, laya-serve exposes the same POST /v1/systemone endpoint and accepts every Jev question shape (choice, score, noul). Switching requires changing only the base URL in your client configuration. Install laya-serve with pip install "laya[serve]", then start it with LAYA_DEVICE=cuda LAYA_PRELOAD=1 laya-serve for GPU inference with all checkpoints preloaded.
What languages does Laya support?
The multilingual checkpoint achieves usable accuracy, defined as above 3 times random baseline, on 45 of the 51 languages tested on the MASSIVE benchmark. The Router detects script and language automatically in under 0.5 ms and dispatches to the appropriate checkpoint, so no manual language configuration is needed. The English checkpoint is optimized for English only; sending non-Latin scripts to it produces unreliable results regardless of confidence score.
Does Laya work without a GPU?
Yes, Laya runs on CPU with the standard PyTorch build. The ONNX Runtime path (laya[onnx]) adds a CPU-optimized backend, including INT8 quantization via the included export script. CPU inference is slower than GPU: the Router with preload shows 193 to 464 ms per request on CPU, compared to 32.8 ms on a T4 GPU. For latency-sensitive production workloads, GPU hosting is recommended.
Similar open-source tools#
Kev
Train and run Jev-like decision models on your own infrastructure
FckSignups
Open-source tools that work instantly, no signup required
fmt
Fast, type-safe C++ formatting that replaces printf and iostreams
agency-agents
Expert AI agent personalities for every workflow
rakazo
AI teammates you own: your keys, your model, your machine.
paperclip
Self-hosted AI agent management with org charts and budgets

