Architecture cost review
Jev vs Laya: Which Decision Model Costs Less for Your Workload?
Both models answer typed questions instead of writing text, and both return probabilities. Speed, privacy and licence favour Laya. Zero-shot accuracy and context length favour Jev. Cost depends on volume and state size.
Direct answer
Choose Jev for good zero-shot answers, many-option questions and long states up to 32,000 tokens, at about $21 per million 500-token decisions. Choose Laya when decisions must stay on your own hardware, latency must be in the tens of milliseconds, or volume passes roughly 11 million decisions a month. Budget for Laya fine-tuning either way, because its base checkpoints trail a majority-class baseline.
Decision summary
| Decision area | What matters |
|---|---|
| Price | Jev: $0.042/M input, free output · Laya: free weights, you pay hardware |
| Context | Jev: 32,000 tokens · Laya: 512 (English) to 1,024 default (8,192 multilingual max) |
| Latency | Jev: ~0.24 s median round trip on OpenRouter · Laya: ~32.8 ms p50 on a T4 |
| Many options (Banking77, 77 labels) | Jev 0.870 · Laya 0.425 (third-party) |
| Calibration (ECE) | Jev 0.246 · Laya 0.081 after temperature refit, 0.466 before |
| Data path | Jev: state sent to TypeSafe via OpenRouter · Laya: stays on your hardware, can run air-gapped |
The quick verdict
Jev and Laya belong to the same new category, decision or System One models. You send state and questions, and you get typed answers with probabilities. Jev is TypeSafe's hosted model on OpenRouter at $0.042 per million input tokens. Laya is Convai Innovations' Apache-2.0 open-weight answer, released three days later.
For most teams starting out, Jev is cheaper in practice because it is useful zero-shot and needs no infrastructure. Laya becomes the cheaper option at high, steady volume. It is also the only option when data cannot leave your environment.
Cost at three volumes
Assume a 500-token state and OpenRouter's 5.5% Standard credit fee, which makes Jev about $22.16 per million decisions. Assume Laya runs on one always-on T4 at about $252 a month. At 1 million decisions a month, Jev costs about $22 and Laya $252. At 5 million, Jev is about $111. At 20 million, Jev is about $443, and the single-GPU Laya floor is now cheaper by roughly $190 a month.
The picture shifts with state size and availability needs. A 2,000-token state quadruples Jev's cost, so Laya breaks even near 2.8 million decisions. A second GPU for failover doubles Laya's floor. Neither scenario includes the one-off cost of labelling data and fine-tuning Laya, which is usually the largest number in year one.
- 1M decisions/month: Jev ≈ $22 · Laya ≈ $252
- 5M decisions/month: Jev ≈ $111 · Laya ≈ $252
- 20M decisions/month: Jev ≈ $443 · Laya ≈ $252
- All at a 500-token state on one always-on T4 for Laya
Accuracy: where each model breaks
Third-party comparisons show the trade-off clearly. On Banking77, with 77 intent labels, Jev scored 0.870 and Laya 0.425. Laya's options share a fixed token budget, so long option lists crowd out state. On the typed-decisions benchmark, base Laya scored 0.362 against a majority baseline of 0.461, while the fine-tuned checkpoint reached 0.766.
Calibration cuts the other way. After a temperature refit on held-out domain data, Laya's expected calibration error was 0.081, against 0.246 reported for Jev. If your code acts on probability thresholds, a well-calibrated fine-tuned Laya can be easier to tune. You have to do the refit yourself, though.
Latency and where the data goes
Laya's headline is speed: about 32.8 milliseconds median per decision on a T4, roughly seven times faster than the 236 to 276 milliseconds third-party tests recorded for Jev. OpenRouter quotes a 0.24-second median round trip. For a gate inside a user-facing request, or a per-token agent loop, that gap can decide the architecture.
Jev sends your state to TypeSafe through OpenRouter. Laya runs wherever you put it, including an air-gapped network. For regulated data, Laya removes a processor from the data path. With Jev, you must document OpenRouter and TypeSafe in your vendor review.
Context and question design
Jev accepts up to 32,000 tokens of state plus questions. Laya's English checkpoint defaults to 512 tokens, about 320 of them for state, and the other checkpoints to 1,024. A question about a whole conversation or a multi-page document fits Jev directly. Laya needs a summarisation step first, which adds its own model cost and error.
In both cases keep arithmetic, dates and lookups in your own code, and ask about the smallest piece of state that answers the question. That habit is what keeps Jev's bill low and Laya's context window sufficient.
A practical migration path
Start on Jev for every decision point, logging state, question, answer, probability, usage.cost and the eventual correct label. After a few weeks you know which questions are frequent, which are simple, and which need many options or long context.
Move only the frequent, short-state, few-option questions to a fine-tuned Laya, and keep Jev for the rest. That hybrid keeps zero-shot coverage for new questions, puts volume on fixed-cost hardware, and gives you a labelled set to measure both models against every month.
Total cost over the first year
A first-year view changes the answer for many teams. Suppose you run 5 million 500-token decisions a month. Jev costs about $111 a month after the 5.5% credit fee, or roughly $1,330 for the year. Laya on two T4s for availability costs about $504 a month in hardware, or $6,050 a year. Add perhaps $4,000 of labelling and fine-tuning, and Laya's first year costs about $10,000.
At 50 million decisions a month the order reverses. Jev costs about $1,108 a month, or roughly $13,300 a year. Laya's hardware and one-off fine-tuning stay near the same $10,000, because two T4s still have spare capacity at that volume. The crossover sits somewhere in the tens of millions of decisions a month, earlier with long states and later if you need more than two GPUs or a dedicated engineer.
Key takeaways
- →Choose Jev for good zero-shot answers, many-option questions and long states up to 32,000 tokens, at about $21 per million 500-token decisions. Choose Laya when decisions must stay on your own hardware, latency must be in the tens of milliseconds, or volume passes roughly 11 million decisions a month. Budget for Laya fine-tuning either way, because its base checkpoints trail a majority-class baseline.
- →Prototype on Jev to learn which questions matter and to collect labels, then move high-volume or residency-bound questions to a fine-tuned Laya once the numbers justify it.
- →Latency and accuracy figures come from Convai's model card and third-party benchmarks, not our own testing; hardware prices vary widely.
How this page was prepared
This September 2026 cluster uses OpenRouter's live models API, OpenRouter and TypeSafe documentation, and the Laya model cards, all checked on 1 October 2026. Vendor benchmark claims are attributed, third-party benchmarks are labelled as such, and every cost example states its token assumptions. We did not run a private benchmark for these pages.
Frequently asked questions
Is Laya better than Jev?
Laya is faster, open-weight and better calibrated after a temperature refit. Jev is far more accurate zero-shot, handles many-option questions much better, and reads 32,000 tokens of state. The better choice depends on volume, latency and data residency.
Which is cheaper, Jev or Laya?
At a 500-token state, Jev is cheaper below roughly 11 million decisions a month compared with one always-on T4 for Laya, before counting Laya's fine-tuning time. Above that, Laya's fixed cost wins.
Can Laya run without sending data to a third party?
Yes. Laya is Apache-2.0 open weights and can run on your own servers or air-gapped. Jev sends state to TypeSafe through OpenRouter.
Do Jev and Laya use the same question types?
Both answer typed questions with probabilities — pick one option, yes or no, or a score on a scale — so code written around one maps closely to the other, though request formats differ.