Architecture cost review
Laya Pricing: What the Free Decision Model Costs to Run
Open weights move cost from an invoice to your own infrastructure and engineering time. Laya is fast and small, but the bill for making it accurate on your labels is real.
Direct answer
Laya costs nothing to licence: it is Apache-2.0 open weights from Convai Innovations. Running it is not free. One always-on rented T4 is about $252 a month at $0.35 an hour, so Laya beats Jev's metered price only above roughly 11 million 500-token decisions a month. The bigger cost is fine-tuning, because the base checkpoints score below a majority-class baseline zero-shot.
Decision summary
| Decision area | What matters |
|---|---|
| Licence | Apache-2.0 open weights from Convai Innovations, released 18 September 2026 |
| Checkpoints | laya 421M (512 ctx) · laya-multilingual 322M (1,024 ctx) · laya-typed-decisions 421M (1,024 ctx) |
| Latency | p50 about 32.8 ms per decision on a Tesla T4 (Convai) |
| Hardware cost example | One always-on T4 at ~$0.35/hour ≈ $252/month |
| Zero-shot accuracy | Base 0.362 vs majority baseline 0.461 on the typed-decisions benchmark |
| Fine-tuned accuracy | 0.766 for laya-typed-decisions on the same benchmark |
What Laya is and what you are not paying for
Laya is an open-source decision-model family from Convai Innovations, published on Hugging Face on 18 September 2026, three days after TypeSafe launched Jev. Like Jev, it takes a state plus typed questions and returns typed answers with probabilities in a single forward pass. It does not generate text.
There is no per-token price, no seat and no rate card. The weights are Apache-2.0, so commercial use, modification and private deployment are allowed. What you buy instead is compute, storage, monitoring and the engineering time to make the model accurate on your own decisions.
The three checkpoints and their limits
The English laya checkpoint is a 421M-parameter decision head on ModernBERT-large with a 512-token default context. After the option budget, that leaves roughly 320 tokens for state. laya-multilingual uses mmBERT-base at 322M parameters with a 1,024-token default context, extendable to 8,192, and covers more than 100 languages. laya-typed-decisions is the 421M model fine-tuned on a typed-decisions benchmark's training split.
These windows are much smaller than Jev's 32,000 tokens. Laya suits short states such as a ticket subject, a single message or a pending tool call. It is a poor fit for whole email threads or documents, unless you summarise first and pay for that summarisation step.
Hardware cost and capacity
Convai reports a median latency of about 32.8 milliseconds per decision on a Tesla T4, and the model runs in under 1 GB of memory. At that speed one T4 handles roughly 30 sequential decisions a second, about 79 million a month at full utilisation. Capacity is rarely the constraint; the always-on instance is.
At an indicative $0.35 an hour, one T4 running all month costs about $252. Add a second instance for availability and the floor is about $504 a month before monitoring, logging and on-call time. Teams that already run GPU inference can absorb this; teams that do not should price the operational work honestly.
Break-even against Jev
Jev costs $0.042 per million input tokens on OpenRouter, or about $0.0000222 per 500-token decision after the 5.5% credit fee. Divide $252 of monthly hardware by that, and break-even is about 11.4 million decisions a month. With a 2,000-token state, Jev costs four times as much per decision and break-even falls to about 2.8 million a month.
Engineering time moves the line sharply. Forty hours of labelling, fine-tuning and evaluation at $100 an hour is $4,000. That would buy about 180 million Jev decisions at 500 tokens. Below tens of millions of decisions a month, the metered API is usually cheaper once people are counted.
- 500-token state: break-even ≈ 11.4M decisions/month on one T4
- 2,000-token state: break-even ≈ 2.8M decisions/month on one T4
- Two T4s for availability double the hardware floor
- One-off fine-tuning effort often outweighs a year of low-volume Jev usage
The accuracy cost hidden in a free model
On the typed-decisions benchmark of 2,000 decisions across four workflows, the base English checkpoint scored 0.362 and the multilingual one 0.352. Random choice scores 0.318, and always picking the most common label scores 0.461. So the base checkpoints sit below the majority-class baseline zero-shot. Only the fine-tuned laya-typed-decisions checkpoint, at 0.766, is clearly useful there.
Calibration needs work too. Convai reports an expected calibration error of 0.081 after refitting temperature on held-out domain data, against 0.466 out of the box. If your code branches on probability thresholds, budget for a held-out labelled set and a calibration step, or the thresholds will be wrong.
When Laya is the right buy
Laya is the better economic choice when decisions must stay inside your network, when the per-decision latency budget is tens of milliseconds, or when volume is high enough that metered tokens dominate. It also suits edge devices and Apple Silicon, where community ports report very low latency without a GPU.
It is the wrong choice for a small team that needs good zero-shot answers on many-option questions this week. In that case, start on Jev or Solar Decide through OpenRouter, collect labelled decisions, and revisit Laya once you have a dataset to fine-tune on and the volume to justify running it.
Hidden operating costs to budget
Self-hosting adds work that a metered API hides. You need a serving stack, health checks, autoscaling or a fixed capacity plan, model versioning, and a rollback path when a fine-tuned checkpoint misbehaves. Each new label set or question type means another round of labelling, training, calibration and evaluation before it reaches production.
Monitoring matters more for a decision model than for a chat model, because code acts on the answers without a person reading them. Track the distribution of answers and probabilities, compare samples against human labels each week, and alert when either drifts. Budget a few engineer-hours a month for this even when nothing breaks. That ongoing time, not the GPU, is often the line that decides whether Laya is cheaper than Jev for a small team.
Key takeaways
- →Laya costs nothing to licence: it is Apache-2.0 open weights from Convai Innovations. Running it is not free. One always-on rented T4 is about $252 a month at $0.35 an hour, so Laya beats Jev's metered price only above roughly 11 million 500-token decisions a month. The bigger cost is fine-tuning, because the base checkpoints score below a majority-class baseline zero-shot.
- →Choose Laya when you need on-premises or air-gapped decisions, sub-50 ms latency, or very high volume, and budget a labelled fine-tuning set before launch.
- →GPU prices vary by provider and region, and the T4 rate is an indicative third-party figure; Laya's benchmark numbers come from Convai and community tests, not our own runs.
How this page was prepared
This September 2026 cluster uses OpenRouter's live models API, OpenRouter and TypeSafe documentation, and the Laya model cards, all checked on 1 October 2026. Vendor benchmark claims are attributed, third-party benchmarks are labelled as such, and every cost example states its token assumptions. We did not run a private benchmark for these pages.
Frequently asked questions
Is Laya free?
The weights are free under Apache-2.0. You pay for the compute to run them, the engineering to fine-tune and calibrate them, and the operations to keep the service up.
How much does it cost to self-host Laya?
One always-on Tesla T4 at about $0.35 an hour is roughly $252 a month, and about $504 with a second instance for availability. Labelling and fine-tuning time comes on top.
At what volume is Laya cheaper than Jev?
At a 500-token state, one T4 breaks even with Jev at about 11.4 million decisions a month before engineering time. At a 2,000-token state, it breaks even at about 2.8 million.
Can I use the base Laya checkpoint without fine-tuning?
Not reliably. On the typed-decisions benchmark the base checkpoints scored below the majority-class baseline, while the fine-tuned checkpoint reached 0.766 accuracy.