Skip to content
All master reviews

Review Machine · Master Review · GitHub repositories

Laya (NandhaKishorM/laya)

The most-starred new AI repository of September, read against its own tracker

Across eighteen sources read on 29 September 2026, Laya is the most-starred new AI repository of September: a 421M-parameter decision engine that answers one question in 32.8 ms on a Tesla T4 under Apache-2.0, whose published 0.766 accuracy belongs to a checkpoint fine-tuned on the benchmark's own training split, while the base weights score 0.362 zero-shot.

Repository created 18 September 2026Licence Apache-2.0Latest release 0.3.21, 27 September 2026Stars on 29 September 27,682Read 29 September 2026Skip to the verdictRead as text

Verdict age

EarlyEleven days old · no owner has had it long

Everything read here is eleven days old or less. The repository was created on 18 September 2026 and its last push was on 27 September. The independent 78-case test and the four editorial write-ups fall between 20 and 22 September, and the three news reports between 19 and 20 September. Nothing describes a deployment running for a month, the version numbers move several times a day, and the tracker already carries reports the next release may answer.

  • Early · editors only
  • Provisional · editors in, owners thin
  • Settled · owner reviews read over months

From 18 sources to one verdict

Seven levels, in the same order in every Master Review. Each line is the takeaway; open a level for the evidence under it, or stop when you have enough.

What we read

18 sources, 19 Sep – 29 Sep 2026. Editorial, news and other tiers thin.

By tier, heaviest first

T1Official4

The project's README and three model cards, read on 29 September 2026.

T2Editorial4Thin

Four write-ups, 20 to 22 September 2026, one of them an independent 78-case run. None had the model in hand for long.

T2News3Thin

Three reports, 19 and 20 September 2026. All three carry the project's own measurements rather than measuring their own.

T3Buyers2

No buyer platform applies to a software subject. This tier holds the usage signals instead: the GitHub REST API's figures for the repository and PyPI's own counts, one call each on 29 September 2026.

Forum4

The Hacker News launch thread of 19 September 2026, and three reports on the project's own tracker from 21 and 22 September 2026.

Other1Thin

One hands-on write-up of 27 September 2026 that ran the model on a CPU.

Date window

19 Sep 2026 to 29 Sep 2026

Editorial reviews 20 Sep – 22 Sep 2026 · News 19 Sep – 20 Sep 2026Every source was read on 29 September 2026, and no source is dated later than 27 September 2026, so the reading covers the repository's first eleven days.

Read, then rejected

  • TypeSafe AI's launch post for Jev · the rival maker's own page about its own model, a different subject; nothing on it bears on this repository
  • The author's March 2025 preprint, SalesRLAgent · about sales conversion prediction, a different subject; it is the evidence behind a priority claim, not evidence about this model
  • The Open Weights, 'Convai's Laya targets fast AI decisions' · says the licence is a custom other licence and that no figures were published, both of which the model card it cites contradicts

What they measured

One question in 32.8 ms on a Tesla T4, 7.2 ms batched. Zero-shot, the base checkpoints score 0.362 and 0.352 where a guess scores 0.318.

Repository figures, 29 September 2026

27,682starseleven days after the repository was created

GitHub API · read 29 Sep

2,413forks

GitHub API · read 29 Sep

115PRsopen pull requests, against 58 open issue reports

GitHub API · read 29 Sep

One question, on a Tesla T4

GitHub · read 29 Sep

Typed-decisions accuracy, 2,000 decisions

The independent suite, 78 cases

0.590microLaya, last of the four open models

Ap[e]Chat Blog · 20 Sep

0.795microGLiNER2, the next best open model

Ap[e]Chat Blog · 20 Sep

0.885microa 27B chat model at about one bit per weight

Ap[e]Chat Blog · 20 Sep

0.974microJev's hosted API, from the suite author's run

Ap[e]Chat Blog · 20 Sep

Time per case, the same suite

30ms/caseLaya, the fastest in the run

Ap[e]Chat Blog · 20 Sep

302ms/caseJev, hosted

Ap[e]Chat Blog · 20 Sep

Calibration, mean expected error

0.466the English checkpoint as shipped

GitHub · read 29 Sep

0.081after one temperature per question type

GitHub · read 29 Sep

0.314the multilingual checkpoint as shipped, with no fitted temperatures

Hugging Face · read 29 Sep

0.106after refitting

Hugging Face · read 29 Sep

Narrow tasks the project reports

0.993email spam filtering

eesel AI · 22 Sep

0.980phishing detection

eesel AI · 22 Sep

0.522ten-way support routing

eesel AI · 22 Sep

Releases

29releasesnine of them uploaded on 24 September 2026

PyPI · read 29 Sep

Languages clearing three times random, of 51

23langsthe English checkpoint

GitHub · read 29 Sep

45langsrouted to the multilingual checkpoint

GitHub · read 29 Sep

Choice questions at the default budget

0.425Laya on the 77-label Banking77 set

GitHub · read 29 Sep

0.870Jev 1.13.0 on 72 labels, published

GitHub · read 29 Sep

Option order, the tracker's catalog sample

69of 100the multilingual checkpoint, with the options and their descriptions reversed

GitHub · 22 Sep

65of 100the English checkpoint

GitHub · 22 Sep

76of 100the typed-decisions checkpoint

GitHub · 22 Sep

Where they agree

Fast, open and candid about its limits, and the sources agree on all three. What they return to is what the weights cannot do untrained.

One tick per editorial review, in order of publication:AE Ap[e]Chat BlogRW Robot WorldEA eesel AIPR Prolyz

Answers in tens of milliseconds on one modest GPU

3/4

The README attaches the test: one question in 32.8 ms or 39.5 ms on a Tesla T4, and 7.2 ms per question at ten in one pass. T1Official T2Editorial T2News Other

Apache-2.0 weights you can host yourself

2/4

Three checkpoints, no per-token bill, and no round trip for the state. T1Official T2Editorial T2News Other

Documentation that names its own failures

2/4

The README's limits section is the reason five write-ups give for trusting the rest of it. T2Editorial T2News Other

Calibration, after one temperature per question type

3/4

Mean expected error 0.466 to 0.081 on the English checkpoint, and 0.314 to 0.106 on the multilingual one. T1Official T2Editorial T2News

The narrow decisions it is built for

1/4

Email spam filtering 0.993, phishing detection 0.980, and ten-way support routing 0.522. T1Official T2Editorial T2News

Attention, not quality

GitHub's own figures for the repository and PyPI's own count for the package, one call each on 29 September 2026. They measure how many people looked or installed, and nothing about accuracy. T3Buyers

27682stars2413forks109457downloads in the previous month

GitHub's and PyPI's own counts, from one call each. Not reviews we read.

What keeps coming up against it

The base checkpoints, the confidence they ship with, wide label sets, and the multilingual path.

The 0.766 belongs to a fine-tuned checkpoint

Well documented · stated by the project first
Documented byT1The projectT1The cardT2EditorialT2NewsT3GitHubForumHow commonthe README and the base model card state it, and seven of the write-ups we read repeat it
T1The project

The README and the base model card give the zero-shot figures beside the fine-tuned one, and call Laya a fast base to specialise rather than a zero-shot decision engine.

GitHub · read 29 SepHugging Face · read 29 Sep

T1The card

The typed-decisions card says the 0.766 comes from a checkpoint fine-tuned on that benchmark's own 1,200-case training split, and that it is not recommended outside its four workflows.

Hugging Face · read 29 Sep

T2Editorial

2/4Both write-ups that lead on accuracy carry the zero-shot figure beside the fine-tuned one, and both tell the reader to plan for fine-tuning.

T2News

Both reports put the caveat in the same paragraph as the headline number.

Tech AI Wire · 19 SepAI Weekly · 19 Sep

T3GitHub

The repository's own figures describe attention, not accuracy: 27,682 stars and 2,413 forks in eleven days.

GitHub API · read 29 Sep

Forum

The most-repeated criticism in the launch thread was that requiring a fine-tune puts Laya in a different category from a model sold as usable untrained.

Hacker News · 19 Sep

Zero-shot, the weights you download are near chance

2/4

The base checkpoints score 0.362 and 0.352 on 2,000 typed decisions, against 0.318 for a random guess and 0.461 for the majority class. T1Official T2Editorial T2News Other

Both checkpoints ship over-confident

3/4

The multilingual one ships with no fitted temperatures at all, so its probabilities read sharper than they are. T1Official T2Editorial T2News

Choice questions are the weak primitive

2/4

Past about twenty options the shared token budget leaves three or four tokens a label, and the order of the options moves the answer. T1Official T2Editorial Forum

The multilingual path needs the router, and the router has its own faults

2/4

The English checkpoint collapses outside English and stays confident while it does, and the router's default holds one checkpoint and sends some Latin-script text to the English one. T1Official T2Editorial Forum

What the usage signals do not say

Not counted toward any complaint here: a star or a download is not a report about the model. T3Buyers Not evidence of quality

GitHub's and PyPI's own counts, from one call each.

Where they split

Four places the evidence pulls two ways, starting with whether 0.766 and 0.974 are answering the same question.

Speed against accuracy, in two rooms

T1Official

The project's own table puts the fine-tuned checkpoint at 0.766 across 2,000 decisions, above Jev's published 0.727 and above the 0.735 ceiling of the teacher it was trained to copy.

T2Editorial

Against that: An independent run of 78 cases on a Mac mini put Laya last of four open models at 0.590 micro accuracy, with Jev's hosted API at 0.974, a one-bit 27B at 0.885 and GLiNER2 at 0.795.

Both honestOne suite is the benchmark the fine-tuned checkpoint trained on, and the other is 78 cases. Speed held in both.

Whether 7.8 times means anything

T1Official

The README sets its 32.8 ms against Jev's third-party 236 to 276 ms and calls it 7.8 times faster, and states plainly that it never measured Jev.

T2Editorial

Against that: Jev's published 70 to 500 ms is an end-to-end figure over the network, while Laya's is one local forward pass. The two are not the same measurement.

SoA local loop against a hosted round trip. The latency that settles it is the one measured on your own workload.

Whether the idea was first

T2News

It reports the author's claim that he published a non-autoregressive, reinforcement-trained decision framework in a March 2025 preprint, a year before TypeSafe launched Jev without papers, open weights or datasets.

Forum

Against that: Commenters in the 318-comment launch thread said the concept has academic precursors, that GLiNER is a closer published match, and that shipping a usable product was the funded lab's contribution.

Not ours to settleNothing read here shows Jev reusing any code. A priority claim is not a finding of copying, and the claim is the author's.

Whether the pace is momentum or churn

T3Buyers

PyPI lists 29 releases in eleven days, nine uploaded on 24 September, and the GitHub API counted 115 open pull requests against 58 open issues, with no push since 27 September.

T1Official

Against that: The release notes for 0.3.21 describe finished work: ONNX parity, an opt-in abstention flag, batch calls on every surface and stricter input checks.

Both trueA repository moving this fast leaves a busy queue and a long changelog at the same time.

Who it's for, who should pass

For a team with labelled decisions of its own, one modest GPU, and a reason to keep the state at home.

It suits you if

  • You have a few thousand labelled decisions of your own and one T4-class GPU.The project's notebook reproduces the fine-tuned checkpoint in four to five hours on two free T4s, and every source reporting 0.766 also reports what the base checkpoints score without that work.
  • You need the decision to stay on your own hardware.Three checkpoints under Apache-2.0, one modest GPU, no per-token bill and no round trip.

Pass if

  • You want a decision engine that works without training.On the project's own benchmark the base checkpoints score 0.362 and 0.352, against 0.318 for a guess and 0.461 for the majority class.
  • Your label set runs past about twenty options, or the order of your options will not hold still.Banking77 is 0.425 against Jev's 0.870 at the default budget, and one tracker report measured the answer changing with the option order.
  • You will trust the confidence number as it ships.Both checkpoints are over-confident as released, and the multilingual one has no fitted temperatures to trust.
  • You need long, noisy documents read reliably.A tracker report found zero-shot decisions on long multilingual posts landing near the majority-class baseline.

The verdict

Fast and open, at a speed nobody disputes. The accuracy is a training run away, and the sources say so themselves.
VerdictEarlyEleven days old · no owner has had it long

A 421M-parameter decision engine that answers one question in tens of milliseconds on a modest GPU, whose best published accuracy belongs to a checkpoint fine-tuned on the benchmark's own split.

Confidence, by tier

Officialstrong

The project's README and three model cards, read on 29 September 2026, state the zero-shot and fine-tuned figures, the calibration gaps and the licence.

Editorialthin

Four write-ups, 20 to 22 September 2026, one of them an independent 78-case run. None had the model in hand for long.

Newsthin

Three reports, 19 and 20 September 2026, all carrying the project's own measurements.

Buyersstrong

The GitHub REST API and PyPI, one call each on 29 September 2026. Counts of attention and installs, not of quality.

Forumstrong

A launch thread with 1,361 points and 318 comments, and three reports on the project's own tracker.

Otherthin

One hands-on write-up that ran the model on a CPU and reported one misclassified ticket in four.

Editorial evidence
under a month old
Newest report
20 Sep 2026
Read
29 Sep · month 0

Share line

It decides in 33 milliseconds. Knowing what to decide is your job.
Review MachineLaya (NandhaKishorM/laya) · Master Review

Rests onThe README's measured 32.8 ms for one question on a Tesla T4, against the base checkpoints' 0.362 zero-shot on the project's own benchmark, where a random guess scores 0.318.

Sources

18 sources, heaviest tier first. Every figure above comes from one of these.

  1. T1OfficialThe project's READMEread 29 Sep
  2. T1Officialconvaiinnovations/laya model cardread 29 Sep
  3. T1Officialconvaiinnovations/laya-multilingual model cardread 29 Sep
  4. T1Officialconvaiinnovations/laya-typed-decisions model cardread 29 Sep
  5. T2EditorialThree Jev clones and a 27B20 Sep
  6. T2EditorialLaya, the Open-Source System 1 Decision Model21 Sep
  7. T2EditorialLaya AI review: the open 33ms decision model, tested honestly22 Sep
  8. T2EditorialJev vs Laya: hosted or open-weight decision models?22 Sep
  9. T2NewsConvai ships Laya, a 421M ModernBERT decision model, Apache 2.019 Sep
  10. T2NewsLaya is a 421M open-weights answer to Jev19 Sep
  11. T2NewsTypeSafe Jev: System One Model, answered by Laya20 Sep
  12. T3BuyersGitHub REST API, repository figuresread 29 Sep
  13. T3Buyerslaya on PyPIread 29 Sep
  14. ForumI built non-autoregressive decision models with RL a year ago19 Sep
  15. ForumIssue 99, zero-shot on long noisy multilingual posts21 Sep
  16. ForumIssue 172, the router's default22 Sep
  17. ForumIssue 171, catalog selection and presentation sensitivity22 Sep
  18. OtherLaya: replace your LLM-as-a-judge with a 322M-parameter decision engine27 Sep

MethodOn 29 September 2026 we read the project's README and three model cards, four editorial write-ups, three news reports, a launch thread with 318 comments and three reports on the project's own issue tracker, and the GitHub and PyPI APIs, by web search and plain page fetch, and synthesised them with AI; nothing was installed, prompted or tested.

Dates without a year are 2026.

The review as text

The same review as one piece of writing · 7 min read

Laya is an open decision engine published on GitHub by NandhaKishorM on 18 September 2026. It reads a piece of text and a set of typed questions and answers with a label, a score and yes or no probabilities, from one forward pass of a 421M-parameter encoder, generating no text at all. Eleven days on, GitHub's REST API counted 27,682 stars and 2,413 forks, the highest of any AI repository created that month. Stars are attention, not quality.

We read the project's README and three model cards, four editorial write-ups, three news reports from 19 to 20 September 2026, its Hacker News launch thread and three reports on its own tracker, and the GitHub and PyPI APIs, all on 29 September 2026. Nothing was installed, prompted or tested.

Consensus

The eighteen sources agree on three things. Laya is quick: one question in tens of milliseconds on a single T4-class GPU. Laya is open: three checkpoints under Apache-2.0, no per-token bill. And as shipped it is not a decision engine anyone can use untrained. The project says so first, in its README and both base model cards.

That is an odd shape for a release of 27,682 stars: the claim it went viral on is one its own documentation qualifies at length. Confidence is moderate. The documentation is dated and detailed, but two write-ups work from the project's own figures, the one independent test we found runs 78 cases, and no source describes a deployment older than a few weeks.

Recurring strengths

Speed, measured and repeated. Ten of the eighteen sources state the latency figure, and the README attaches the test: 32.8 ms for one question on the multilingual checkpoint, 39.5 ms on the English one, on a Tesla T4, falling to 7.2 ms and 15.9 ms per question at ten in one pass. The independent test that put Laya last on accuracy still timed it first, at 30 ms a case against about 302 ms for Jev.

An open licence and no meter. Seven sources name it: Apache-2.0, three checkpoints, weights on Hugging Face, a package on PyPI. For a team whose state cannot leave its own network, that is the whole argument.

Documentation that names its own failures. Five write-ups single out the README's limits section as the reason to trust the rest of it.

Calibration, once somebody refits the temperatures. Six sources report it. Both checkpoints ship over-confident, and refitting one temperature per question type moves the English checkpoint's mean expected calibration error from 0.466 to 0.081 and the multilingual one's from 0.314 to 0.106.

The narrow binary decisions. Spam filtering at 0.993 and phishing detection at 0.980 in the review we read, with a ten-way routing task at 0.522.

Recurring complaints

Zero-shot, the weights you download are near chance. Nine sources state it and the project's own table is the plainest. On 2,000 typed decisions the English base scores 0.362 and the multilingual base 0.352, against 0.318 for a random guess and 0.461 for always answering with the plurality label. The 0.766 in the headlines belongs to a third checkpoint, fine-tuned on that benchmark's own training split.

Both checkpoints ship over-confident. Six sources raise it; the multilingual one ships with no fitted temperatures at all. A hands-on write-up that ran the model on plain CPU reported seconds per ticket there rather than milliseconds, and one misclassified ticket in four.

Choice questions are the weak primitive. Four sources and the tracker line up. Past about twenty options accuracy falls away, because every option shares a fixed token budget: on the 77-label Banking77 set that is three or four tokens a label and 0.425 against Jev's 0.870. Order matters too: one tracker report found reversing the options and their descriptions changed the selection in 69 of 100 catalog cases on the multilingual checkpoint, 65 on the English one and 76 on the typed one.

The multilingual path is where the seams show. Four sources and one tracker report. The English checkpoint collapses outside English and stays confident while it does: over a 51-language sweep it averages 0.227 with mean calibration error 0.733, only 23 of the 51 clear three times random, and Khmer scores 0.000 at 0.952 confidence. Routed, 45 of the 51 clear it. The router has its own report: its default holds one checkpoint, reloading on every switch, and Spanish and Italian route to the English checkpoint.

Where reviewers split

Speed against accuracy, in two rooms. The project's table puts the fine-tuned checkpoint at 0.766 across 2,000 decisions, above Jev's published 0.727 and above the 0.735 teacher ceiling. The independent 78-case run on a Mac mini put Laya last of four open models at 0.590, behind Jev's hosted API at 0.974, a one-bit 27B at 0.885 and GLiNER2 at 0.795. Both are honest and answer different questions: one suite is the benchmark the checkpoint trained on. Speed held in both.

Whether 7.8 times means anything. The README sets its 32.8 ms against Jev's third-party 236 to 276 ms and calls it 7.8 times faster, saying plainly that it never measured Jev. A comparison piece we read notes that Jev's 70 to 500 ms is an end-to-end figure over the network while Laya's is one local forward pass.

Whether the idea was first. The news report we read describes the author's claim that he published a non-autoregressive, reinforcement-trained decision framework in a March 2025 preprint, a year before TypeSafe launched Jev without papers, open weights or datasets. Commenters in the 318-comment launch thread answered differently: that the concept has academic precursors, that GLiNER is a closer published match, and that the funded lab's contribution was a usable product. Nothing read here shows Jev reusing any code, and a priority claim is not a finding of copying.

Whether the pace is momentum or churn. PyPI lists 29 releases in eleven days, nine uploaded on 24 September and 109,457 downloads in the previous month, and the GitHub API counted 115 open pull requests against 58 open issues, with no push since 27 September. The README's release notes for 0.3.21 describe finished work. Both are true: a repository this fast leaves a busy queue and a long changelog.

Who it suits

You have a few thousand labelled decisions of your own and one T4-class GPU. The project's notebook reproduces the fine-tuned checkpoint in four to five hours on two free T4s, and every source reporting 0.766 also reports what the base checkpoints score without that work.

You need the decision to stay on your own hardware. Three checkpoints under Apache-2.0, one modest GPU, no per-token bill and no round trip.

Who should pass

You want a decision engine that works without training. On the project's own benchmark the base checkpoints score 0.362 and 0.352, against 0.318 for a guess.

Your label set runs past about twenty options, or the order of your options will not hold still. Banking77 is 0.425 against Jev's 0.870 at the default budget, and one tracker report measured the answer changing with the option order.

You will trust the confidence number as it ships. Both checkpoints are over-confident as released, and the multilingual one has no fitted temperatures to trust.

You need long, noisy documents read reliably. A tracker report found zero-shot decisions on long multilingual posts landing near the majority-class baseline.

One verdict a week: the most useful Master Review we finished, the complaint that kept appearing, and who should skip it.

Sources

On 29 September 2026 we read the project's README and three model cards, four editorial write-ups, three news reports, a launch thread with 318 comments and three reports on the project's own issue tracker, and the GitHub and PyPI APIs, by web search and plain page fetch, and synthesised them with AI; nothing was installed, prompted or tested.

One verdict a week.

Every week, the most useful Master Review we finished: what the internet agrees on, the complaint that kept appearing, and who should skip it.

By subscribing you agree to receive one weekly email from Review Machine. You can unsubscribe at any time.