Review Machine · Master Review · GitHub repositories
Laya (NandhaKishorM/laya)
The most-starred new AI repository of September, read against its own tracker
Across eighteen sources read on 29 September 2026, Laya is the most-starred new AI repository of September: a 421M-parameter decision engine that answers one question in 32.8 ms on a Tesla T4 under Apache-2.0, whose published 0.766 accuracy belongs to a checkpoint fine-tuned on the benchmark's own training split, while the base weights score 0.362 zero-shot.
Repository created 18 September 2026Licence Apache-2.0Latest release 0.3.21, 27 September 2026Stars on 29 September 27,682Read 29 September 2026Skip to the verdictRead as text

Verdict age
EarlyEleven days old · no owner has had it longEverything read here is eleven days old or less. The repository was created on 18 September 2026 and its last push was on 27 September. The independent 78-case test and the four editorial write-ups fall between 20 and 22 September, and the three news reports between 19 and 20 September. Nothing describes a deployment running for a month, the version numbers move several times a day, and the tracker already carries reports the next release may answer.
- Early · editors only
- Provisional · editors in, owners thin
- Settled · owner reviews read over months
From 18 sources to one verdict
Seven levels, in the same order in every Master Review. Each line is the takeaway; open a level for the evidence under it, or stop when you have enough.
What we read
18 sources, 19 Sep – 29 Sep 2026. Editorial, news and other tiers thin.
By tier, heaviest first
The project's README and three model cards, read on 29 September 2026.
Four write-ups, 20 to 22 September 2026, one of them an independent 78-case run. None had the model in hand for long.
Three reports, 19 and 20 September 2026. All three carry the project's own measurements rather than measuring their own.
No buyer platform applies to a software subject. This tier holds the usage signals instead: the GitHub REST API's figures for the repository and PyPI's own counts, one call each on 29 September 2026.
The Hacker News launch thread of 19 September 2026, and three reports on the project's own tracker from 21 and 22 September 2026.
One hands-on write-up of 27 September 2026 that ran the model on a CPU.
Date window
19 Sep 2026 to 29 Sep 2026
Editorial reviews 20 Sep – 22 Sep 2026 · News 19 Sep – 20 Sep 2026Every source was read on 29 September 2026, and no source is dated later than 27 September 2026, so the reading covers the repository's first eleven days.
Read, then rejected
- TypeSafe AI's launch post for Jev · the rival maker's own page about its own model, a different subject; nothing on it bears on this repository
- The author's March 2025 preprint, SalesRLAgent · about sales conversion prediction, a different subject; it is the evidence behind a priority claim, not evidence about this model
- The Open Weights, 'Convai's Laya targets fast AI decisions' · says the licence is a custom other licence and that no figures were published, both of which the model card it cites contradicts
What they measured
One question in 32.8 ms on a Tesla T4, 7.2 ms batched. Zero-shot, the base checkpoints score 0.362 and 0.352 where a guess scores 0.318.
Repository figures, 29 September 2026
27,682starseleven days after the repository was created
GitHub API · read 29 Sep2,413forks
GitHub API · read 29 Sep115PRsopen pull requests, against 58 open issue reports
GitHub API · read 29 SepOne question, on a Tesla T4
GitHub · read 29 SepTyped-decisions accuracy, 2,000 decisions
The independent suite, 78 cases
0.590microLaya, last of the four open models
Ap[e]Chat Blog · 20 Sep0.795microGLiNER2, the next best open model
Ap[e]Chat Blog · 20 Sep0.885microa 27B chat model at about one bit per weight
Ap[e]Chat Blog · 20 Sep0.974microJev's hosted API, from the suite author's run
Ap[e]Chat Blog · 20 SepTime per case, the same suite
30ms/caseLaya, the fastest in the run
Ap[e]Chat Blog · 20 Sep302ms/caseJev, hosted
Ap[e]Chat Blog · 20 SepCalibration, mean expected error
0.466the English checkpoint as shipped
GitHub · read 29 Sep0.081after one temperature per question type
GitHub · read 29 Sep0.314the multilingual checkpoint as shipped, with no fitted temperatures
Hugging Face · read 29 Sep0.106after refitting
Hugging Face · read 29 SepNarrow tasks the project reports
0.993email spam filtering
eesel AI · 22 Sep0.980phishing detection
eesel AI · 22 Sep0.522ten-way support routing
eesel AI · 22 SepReleases
29releasesnine of them uploaded on 24 September 2026
PyPI · read 29 SepLanguages clearing three times random, of 51
23langsthe English checkpoint
GitHub · read 29 Sep45langsrouted to the multilingual checkpoint
GitHub · read 29 SepChoice questions at the default budget
0.425Laya on the 77-label Banking77 set
GitHub · read 29 Sep0.870Jev 1.13.0 on 72 labels, published
GitHub · read 29 SepOption order, the tracker's catalog sample
69of 100the multilingual checkpoint, with the options and their descriptions reversed
GitHub · 22 Sep65of 100the English checkpoint
GitHub · 22 Sep76of 100the typed-decisions checkpoint
GitHub · 22 SepWhere they agree
Fast, open and candid about its limits, and the sources agree on all three. What they return to is what the weights cannot do untrained.
One tick per editorial review, in order of publication:AE Ap[e]Chat BlogRW Robot WorldEA eesel AIPR Prolyz
Answers in tens of milliseconds on one modest GPU
3/4The README attaches the test: one question in 32.8 ms or 39.5 ms on a Tesla T4, and 7.2 ms per question at ten in one pass. T1Official T2Editorial T2News Other
Apache-2.0 weights you can host yourself
2/4Three checkpoints, no per-token bill, and no round trip for the state. T1Official T2Editorial T2News Other
Documentation that names its own failures
2/4The README's limits section is the reason five write-ups give for trusting the rest of it. T2Editorial T2News Other
Calibration, after one temperature per question type
3/4Mean expected error 0.466 to 0.081 on the English checkpoint, and 0.314 to 0.106 on the multilingual one. T1Official T2Editorial T2News
The narrow decisions it is built for
1/4Email spam filtering 0.993, phishing detection 0.980, and ten-way support routing 0.522. T1Official T2Editorial T2News
Attention, not quality
GitHub's own figures for the repository and PyPI's own count for the package, one call each on 29 September 2026. They measure how many people looked or installed, and nothing about accuracy. T3Buyers
27682stars2413forks109457downloads in the previous month
GitHub's and PyPI's own counts, from one call each. Not reviews we read.
What keeps coming up against it
The base checkpoints, the confidence they ship with, wide label sets, and the multilingual path.
The 0.766 belongs to a fine-tuned checkpoint
Well documented · stated by the project firstThe README and the base model card give the zero-shot figures beside the fine-tuned one, and call Laya a fast base to specialise rather than a zero-shot decision engine.
GitHub · read 29 SepHugging Face · read 29 Sep
The typed-decisions card says the 0.766 comes from a checkpoint fine-tuned on that benchmark's own 1,200-case training split, and that it is not recommended outside its four workflows.
Hugging Face · read 29 Sep
2/4Both write-ups that lead on accuracy carry the zero-shot figure beside the fine-tuned one, and both tell the reader to plan for fine-tuning.
Both reports put the caveat in the same paragraph as the headline number.
Tech AI Wire · 19 SepAI Weekly · 19 Sep
The repository's own figures describe attention, not accuracy: 27,682 stars and 2,413 forks in eleven days.
GitHub API · read 29 Sep
The most-repeated criticism in the launch thread was that requiring a fine-tune puts Laya in a different category from a model sold as usable untrained.
Hacker News · 19 Sep
Zero-shot, the weights you download are near chance
2/4The base checkpoints score 0.362 and 0.352 on 2,000 typed decisions, against 0.318 for a random guess and 0.461 for the majority class. T1Official T2Editorial T2News Other
Both checkpoints ship over-confident
3/4The multilingual one ships with no fitted temperatures at all, so its probabilities read sharper than they are. T1Official T2Editorial T2News
Choice questions are the weak primitive
2/4Past about twenty options the shared token budget leaves three or four tokens a label, and the order of the options moves the answer. T1Official T2Editorial Forum
The multilingual path needs the router, and the router has its own faults
2/4The English checkpoint collapses outside English and stays confident while it does, and the router's default holds one checkpoint and sends some Latin-script text to the English one. T1Official T2Editorial Forum
What the usage signals do not say
Not counted toward any complaint here: a star or a download is not a report about the model. T3Buyers Not evidence of quality
GitHub's and PyPI's own counts, from one call each.
Where they split
Four places the evidence pulls two ways, starting with whether 0.766 and 0.974 are answering the same question.
Speed against accuracy, in two rooms
The project's own table puts the fine-tuned checkpoint at 0.766 across 2,000 decisions, above Jev's published 0.727 and above the 0.735 ceiling of the teacher it was trained to copy.
Against that: An independent run of 78 cases on a Mac mini put Laya last of four open models at 0.590 micro accuracy, with Jev's hosted API at 0.974, a one-bit 27B at 0.885 and GLiNER2 at 0.795.
Both honestOne suite is the benchmark the fine-tuned checkpoint trained on, and the other is 78 cases. Speed held in both.
Whether 7.8 times means anything
The README sets its 32.8 ms against Jev's third-party 236 to 276 ms and calls it 7.8 times faster, and states plainly that it never measured Jev.
Against that: Jev's published 70 to 500 ms is an end-to-end figure over the network, while Laya's is one local forward pass. The two are not the same measurement.
SoA local loop against a hosted round trip. The latency that settles it is the one measured on your own workload.
Whether the idea was first
It reports the author's claim that he published a non-autoregressive, reinforcement-trained decision framework in a March 2025 preprint, a year before TypeSafe launched Jev without papers, open weights or datasets.
Against that: Commenters in the 318-comment launch thread said the concept has academic precursors, that GLiNER is a closer published match, and that shipping a usable product was the funded lab's contribution.
Not ours to settleNothing read here shows Jev reusing any code. A priority claim is not a finding of copying, and the claim is the author's.
Whether the pace is momentum or churn
PyPI lists 29 releases in eleven days, nine uploaded on 24 September, and the GitHub API counted 115 open pull requests against 58 open issues, with no push since 27 September.
Against that: The release notes for 0.3.21 describe finished work: ONNX parity, an opt-in abstention flag, batch calls on every surface and stricter input checks.
Both trueA repository moving this fast leaves a busy queue and a long changelog at the same time.
Who it's for, who should pass
For a team with labelled decisions of its own, one modest GPU, and a reason to keep the state at home.
It suits you if
- You have a few thousand labelled decisions of your own and one T4-class GPU.The project's notebook reproduces the fine-tuned checkpoint in four to five hours on two free T4s, and every source reporting 0.766 also reports what the base checkpoints score without that work.
- You need the decision to stay on your own hardware.Three checkpoints under Apache-2.0, one modest GPU, no per-token bill and no round trip.
Pass if
- You want a decision engine that works without training.On the project's own benchmark the base checkpoints score 0.362 and 0.352, against 0.318 for a guess and 0.461 for the majority class.
- Your label set runs past about twenty options, or the order of your options will not hold still.Banking77 is 0.425 against Jev's 0.870 at the default budget, and one tracker report measured the answer changing with the option order.
- You will trust the confidence number as it ships.Both checkpoints are over-confident as released, and the multilingual one has no fitted temperatures to trust.
- You need long, noisy documents read reliably.A tracker report found zero-shot decisions on long multilingual posts landing near the majority-class baseline.
The verdict
Fast and open, at a speed nobody disputes. The accuracy is a training run away, and the sources say so themselves.
A 421M-parameter decision engine that answers one question in tens of milliseconds on a modest GPU, whose best published accuracy belongs to a checkpoint fine-tuned on the benchmark's own split.
Confidence, by tier
The project's README and three model cards, read on 29 September 2026, state the zero-shot and fine-tuned figures, the calibration gaps and the licence.
Four write-ups, 20 to 22 September 2026, one of them an independent 78-case run. None had the model in hand for long.
Three reports, 19 and 20 September 2026, all carrying the project's own measurements.
The GitHub REST API and PyPI, one call each on 29 September 2026. Counts of attention and installs, not of quality.
A launch thread with 1,361 points and 318 comments, and three reports on the project's own tracker.
One hands-on write-up that ran the model on a CPU and reported one misclassified ticket in four.
- Editorial evidence
- under a month old
- Newest report
- 20 Sep 2026
- Read
- 29 Sep · month 0
It decides in 33 milliseconds. Knowing what to decide is your job.
Rests onThe README's measured 32.8 ms for one question on a Tesla T4, against the base checkpoints' 0.362 zero-shot on the project's own benchmark, where a random guess scores 0.318.
Sources
18 sources, heaviest tier first. Every figure above comes from one of these.
- T1OfficialThe project's READMEread 29 Sep
- T1Officialconvaiinnovations/laya model cardread 29 Sep
- T1Officialconvaiinnovations/laya-multilingual model cardread 29 Sep
- T1Officialconvaiinnovations/laya-typed-decisions model cardread 29 Sep
- T2EditorialThree Jev clones and a 27B20 Sep
- T2EditorialLaya, the Open-Source System 1 Decision Model21 Sep
- T2EditorialLaya AI review: the open 33ms decision model, tested honestly22 Sep
- T2EditorialJev vs Laya: hosted or open-weight decision models?22 Sep
- T2NewsConvai ships Laya, a 421M ModernBERT decision model, Apache 2.019 Sep
- T2NewsLaya is a 421M open-weights answer to Jev19 Sep
- T2NewsTypeSafe Jev: System One Model, answered by Laya20 Sep
- T3BuyersGitHub REST API, repository figuresread 29 Sep
- T3Buyerslaya on PyPIread 29 Sep
- ForumI built non-autoregressive decision models with RL a year ago19 Sep
- ForumIssue 99, zero-shot on long noisy multilingual posts21 Sep
- ForumIssue 172, the router's default22 Sep
- ForumIssue 171, catalog selection and presentation sensitivity22 Sep
- OtherLaya: replace your LLM-as-a-judge with a 322M-parameter decision engine27 Sep
MethodOn 29 September 2026 we read the project's README and three model cards, four editorial write-ups, three news reports, a launch thread with 318 comments and three reports on the project's own issue tracker, and the GitHub and PyPI APIs, by web search and plain page fetch, and synthesised them with AI; nothing was installed, prompted or tested.
Dates without a year are 2026.
The review as text
Laya is an open decision engine published on GitHub by NandhaKishorM on 18 September 2026. It reads a piece of text and a set of typed questions and answers with a label, a score and yes or no probabilities, from one forward pass of a 421M-parameter encoder, generating no text at all. Eleven days on, GitHub's REST API counted 27,682 stars and 2,413 forks, the highest of any AI repository created that month. Stars are attention, not quality.
We read the project's README and three model cards, four editorial write-ups, three news reports from 19 to 20 September 2026, its Hacker News launch thread and three reports on its own tracker, and the GitHub and PyPI APIs, all on 29 September 2026. Nothing was installed, prompted or tested.
Consensus
The eighteen sources agree on three things. Laya is quick: one question in tens of milliseconds on a single T4-class GPU. Laya is open: three checkpoints under Apache-2.0, no per-token bill. And as shipped it is not a decision engine anyone can use untrained. The project says so first, in its README and both base model cards.
That is an odd shape for a release of 27,682 stars: the claim it went viral on is one its own documentation qualifies at length. Confidence is moderate. The documentation is dated and detailed, but two write-ups work from the project's own figures, the one independent test we found runs 78 cases, and no source describes a deployment older than a few weeks.
Recurring strengths
Speed, measured and repeated. Ten of the eighteen sources state the latency figure, and the README attaches the test: 32.8 ms for one question on the multilingual checkpoint, 39.5 ms on the English one, on a Tesla T4, falling to 7.2 ms and 15.9 ms per question at ten in one pass. The independent test that put Laya last on accuracy still timed it first, at 30 ms a case against about 302 ms for Jev.
An open licence and no meter. Seven sources name it: Apache-2.0, three checkpoints, weights on Hugging Face, a package on PyPI. For a team whose state cannot leave its own network, that is the whole argument.
Documentation that names its own failures. Five write-ups single out the README's limits section as the reason to trust the rest of it.
Calibration, once somebody refits the temperatures. Six sources report it. Both checkpoints ship over-confident, and refitting one temperature per question type moves the English checkpoint's mean expected calibration error from 0.466 to 0.081 and the multilingual one's from 0.314 to 0.106.
The narrow binary decisions. Spam filtering at 0.993 and phishing detection at 0.980 in the review we read, with a ten-way routing task at 0.522.
Recurring complaints
Zero-shot, the weights you download are near chance. Nine sources state it and the project's own table is the plainest. On 2,000 typed decisions the English base scores 0.362 and the multilingual base 0.352, against 0.318 for a random guess and 0.461 for always answering with the plurality label. The 0.766 in the headlines belongs to a third checkpoint, fine-tuned on that benchmark's own training split.
Both checkpoints ship over-confident. Six sources raise it; the multilingual one ships with no fitted temperatures at all. A hands-on write-up that ran the model on plain CPU reported seconds per ticket there rather than milliseconds, and one misclassified ticket in four.
Choice questions are the weak primitive. Four sources and the tracker line up. Past about twenty options accuracy falls away, because every option shares a fixed token budget: on the 77-label Banking77 set that is three or four tokens a label and 0.425 against Jev's 0.870. Order matters too: one tracker report found reversing the options and their descriptions changed the selection in 69 of 100 catalog cases on the multilingual checkpoint, 65 on the English one and 76 on the typed one.
The multilingual path is where the seams show. Four sources and one tracker report. The English checkpoint collapses outside English and stays confident while it does: over a 51-language sweep it averages 0.227 with mean calibration error 0.733, only 23 of the 51 clear three times random, and Khmer scores 0.000 at 0.952 confidence. Routed, 45 of the 51 clear it. The router has its own report: its default holds one checkpoint, reloading on every switch, and Spanish and Italian route to the English checkpoint.
Where reviewers split
Speed against accuracy, in two rooms. The project's table puts the fine-tuned checkpoint at 0.766 across 2,000 decisions, above Jev's published 0.727 and above the 0.735 teacher ceiling. The independent 78-case run on a Mac mini put Laya last of four open models at 0.590, behind Jev's hosted API at 0.974, a one-bit 27B at 0.885 and GLiNER2 at 0.795. Both are honest and answer different questions: one suite is the benchmark the checkpoint trained on. Speed held in both.
Whether 7.8 times means anything. The README sets its 32.8 ms against Jev's third-party 236 to 276 ms and calls it 7.8 times faster, saying plainly that it never measured Jev. A comparison piece we read notes that Jev's 70 to 500 ms is an end-to-end figure over the network while Laya's is one local forward pass.
Whether the idea was first. The news report we read describes the author's claim that he published a non-autoregressive, reinforcement-trained decision framework in a March 2025 preprint, a year before TypeSafe launched Jev without papers, open weights or datasets. Commenters in the 318-comment launch thread answered differently: that the concept has academic precursors, that GLiNER is a closer published match, and that the funded lab's contribution was a usable product. Nothing read here shows Jev reusing any code, and a priority claim is not a finding of copying.
Whether the pace is momentum or churn. PyPI lists 29 releases in eleven days, nine uploaded on 24 September and 109,457 downloads in the previous month, and the GitHub API counted 115 open pull requests against 58 open issues, with no push since 27 September. The README's release notes for 0.3.21 describe finished work. Both are true: a repository this fast leaves a busy queue and a long changelog.
Who it suits
You have a few thousand labelled decisions of your own and one T4-class GPU. The project's notebook reproduces the fine-tuned checkpoint in four to five hours on two free T4s, and every source reporting 0.766 also reports what the base checkpoints score without that work.
You need the decision to stay on your own hardware. Three checkpoints under Apache-2.0, one modest GPU, no per-token bill and no round trip.
Who should pass
You want a decision engine that works without training. On the project's own benchmark the base checkpoints score 0.362 and 0.352, against 0.318 for a guess.
Your label set runs past about twenty options, or the order of your options will not hold still. Banking77 is 0.425 against Jev's 0.870 at the default budget, and one tracker report measured the answer changing with the option order.
You will trust the confidence number as it ships. Both checkpoints are over-confident as released, and the multilingual one has no fitted temperatures to trust.
You need long, noisy documents read reliably. A tracker report found zero-shot decisions on long multilingual posts landing near the majority-class baseline.
Sources
- GitHub README, official, read 29 September 2026.
- Hugging Face, the laya model card, official, read 29 September 2026.
- Hugging Face, the laya-multilingual model card, official, read 29 September 2026.
- Hugging Face, the laya-typed-decisions model card, official, read 29 September 2026.
- GitHub REST API, repository figures, usage signal, read 29 September 2026.
- PyPI, laya, usage signal, read 29 September 2026.
- eesel AI, Laya AI review, editorial, 22 September 2026.
- Prolyz, Jev vs Laya, editorial, 22 September 2026.
- Robot World, Laya, the open-source System 1 decision model, editorial, 21 September 2026.
- Ap[e]Chat Blog, Three Jev clones and a 27B, editorial, 20 September 2026.
- AI Weekly, Convai ships Laya, news, 19 September 2026.
- Tech AI Wire, an open-weights answer to Jev, news, 19 September 2026.
- Gadget Pilipinas, the System One model, answered by Laya, news, 20 September 2026.
- Hacker News, the launch thread, forum, 19 September 2026.
- GitHub issue 172, the router's default, forum, 22 September 2026.
- GitHub issue 171, catalog selection, forum, 22 September 2026.
- GitHub issue 99, zero-shot on long posts, forum, 21 September 2026.
- dev.to, replace your LLM-as-a-judge, other, 27 September 2026.
On 29 September 2026 we read the project's README and three model cards, four editorial write-ups, three news reports, a launch thread with 318 comments and three reports on the project's own issue tracker, and the GitHub and PyPI APIs, by web search and plain page fetch, and synthesised them with AI; nothing was installed, prompted or tested.
One verdict a week.
Every week, the most useful Master Review we finished: what the internet agrees on, the complaint that kept appearing, and who should skip it.