Review Machine · Master Review · AI models
Llama 4 Scout
What developers use Meta's newest Llama for
Across 14 sources read on 9 October 2026, Llama 4 Scout is described as a cheap reader of documents and images, while its 10-million-token context window is the part the sources dispute.
Checkpoint Llama-4-Scout-17B-16E-InstructSize 17B active, 109B totalReleased 5 April 2025Licence Llama 4 Community LicenseRead 9 October 2026Skip to the verdictRead as text

Verdict age
ProvisionalTwo evaluations · developers thinOfficial pages are from April 2025 and the guides from April to July 2026. A later Llama release, or a new independent long-context evaluation, would change the reading.
- Early · editors only
- Provisional · editors in, owners thin
- Settled · owner reviews read over months
From 14 sources to one verdict
Seven levels, in the same order in every Master Review. Each line is the takeaway; open a level for the evidence under it, or stop when you have enough.
What we read
14 sources, 5 Apr 2025 – 9 Oct 2026. Editorial, other and forum tiers thin. Buyer tier empty.
By tier, heaviest first
Meta's launch post and model card, the Hugging Face page and the licence file.
An evaluation page read 9 Oct 2026 and one April 2025 article. Two pieces only.
11 Apr 2025
No buyer tier exists for a downloadable model.
One launch thread as a partial page view and one API call returning nine comments.
8 Apr – 9 Oct 2026
Date window
5 Apr 2025 to 9 Oct 2026
Editorial reviews 7 Apr 2025 – 9 Oct 2026 · News 11 Apr 2025Launch pages from 5 Apr 2025; guides from 8 Apr 2026 to 13 Jul 2026; pages and figures read 9 Oct 2026.
Couldn’t read, so not used
- BankInfoSecurity articleHTTP 403
- GovInfoSecurity articleHTTP 403
Where a page refused a plain fetch, the source was dropped with every count that rested only on it.
Read, then rejected
- inkeybit Scout guide · gives a month, not a day, so it cannot be dated
- Maverick spec and price pages · a different model from Scout
What they measured
A claimed 10M-token window, 15.6% accuracy at 120,000 tokens in one benchmark, $0.19 per million input tokens.
Context window claimed by Meta
10Mtokensthe maker's claim
Meta · 5 Apr 2025Long-narrative benchmark at 120,000 tokens
15.6%Fiction.Live accuracy as reported in April 2025One benchmark
The Decoder · 7 Apr 2025Hosted price per million tokens
$0.19in
Artificial Analysis · read 9 Oct$0.68out
Artificial Analysis · read 9 OctIntelligence Index
8peer median given as 7Ranking, not use
Artificial Analysis · read 9 OctHugging Face, last month
170,153downloadsNot use
Hugging Face · read 9 Oct1,368likes
Hugging Face · read 9 OctWhere they agree
Four uses recur, and only the price is backed by both evaluation sources.
One tick per editorial review, in order of publication:TD The DecoderAA Artificial Analysis
Reading long documents and whole codebases
0/2Meta's own claim and two independent guides; neither evaluation source names it. T1Official Other
Reading charts, documents and images
0/2T1Official Other
Cheap to call as a hosted model
2/2T2Editorial
Runs on one large GPU or a high-memory desktop once quantised
0/2Speeds are the authors' own reports. T1Official Forum Other
What keeps coming up against it
The long-context window is the complaint that recurs most, in four sources.
The long-context window does not hold up
Documented · size unknownMeta states a 10M-token window and presents whole-codebase work as a use. This is the maker's claim.
Meta · 5 Apr 2025
1/2The Decoder reported 15.6 percent on Fiction.Live narratives of 120,000 tokens.
Two guides doubt the window in practice: one puts reliable use near 32,000 tokens, the other says the single-GPU claim does not show full-window serving.
Zen van Riel · 8 AprVerdent · read 9 Oct
Commenters on the launch thread questioned recall across the whole window.
Hacker News · 5 Apr 2025
Long context fails well before 10M tokens
1/2T2Editorial Forum Other
The memory bill for local use
0/2Forum Other
Weak for coding
0/2One guide's benchmark figure and one developer's comment. Forum Other
Where they split
Sources split on whether the long window is a strength.
Is the 10M-token window a reason to pick Scout?
Presented as the reason to pick it, for codebases and many documents in one pass.
Against that: Reported failing well before 10 million tokens, with a benchmark figure.
Our readThe window is what the model accepts. Whether answers stay reliable is what the benchmark questions.
General ranking against coding
Ranks Scout above its peer median on its index.
Against that: Calls it unsuitable for coding on one benchmark.
Both trueDifferent tests of different things.
A different model: Maverick
A tuned experimental Maverick is shown on a public chat leaderboard.
Against that: The public Maverick ranked 32nd, below several older rivals. Not Scout, and not counted.
Who it's for, who should pass
Suits bulk reading of documents and images; not long-recall or coding work.
It suits you if
- You have documents, charts or screenshots to read in bulk and pay per token.Price and image reading recur in the sources.
- You have a machine with 48 GB or more of memory and want open weights.Two guides and the model card give these requirements.
- You want a long input but will check answers past about 32,000 tokens.The sources disagree on how far recall holds.
Pass if
- You need dependable recall across hundreds of thousands of tokens.The Decoder and two guides report it failing.
- Your main work is writing code.A guide and a developer comment both fault it.
- You want a model that is still being developed.Two reference pages say Meta moved to Muse Spark in April 2026.
The verdict
Used for reading and price, disputed on its headline feature.
By the sources we read, Scout is a cheap reader of documents and images whose 10-million-token window is the part they dispute, from a small sample.
Confidence, by tier
Size, date and claims come straight from Meta's pages.
Two pieces, one from launch week.
Three guides by individual authors.
A partial thread view and nine comments.
No buyer tier exists.
- Editorial evidence
- up to 18 months old
- Newest report
- 11 Apr 2025
- Read
- 9 Oct · month 18
Ten million tokens of context, and a benchmark found trouble at 120,000.
Rests onThe Decoder, April 2025: Scout reached 15.6 percent on Fiction.Live narratives of 120,000 tokens, against Meta's claimed 10M-token window.
Sources
14 sources, heaviest tier first. Every figure above comes from one of these.
- T1OfficialMeta, Llama 4 launch post5 Apr 2025
- T1OfficialLlama 4 Community License Agreementupdated 5 Apr 2025
- T1OfficialMeta, Llama 4 model cardread 9 Oct
- T1OfficialHugging Face, Llama-4-Scout-17B-16E-Instructone page load, headline figures onlyread 9 Oct
- T2EditorialThe Decoder, Llama 4 on standard tests and long context7 Apr 2025
- T2EditorialArtificial Analysis, Llama 4 Scoutevaluation page, one loadread 9 Oct
- T2NewsTechCrunch, vanilla Maverick ranks below rivals11 Apr 2025
- ForumHacker News, The Llama 4 herdpartial view of an ordinary page5 Apr 2025
- ForumHacker News search, comments on Llama 4 Scoutone API call, nine comments shownread 9 Oct
- OtherZen van Riel, Llama 4 Scout practical guide8 Apr
- OtherDEV Community, Llama 4 Scout locally13 Jul
- OtherMungomash, Llama versionsupdated 29 Sep
- OtherVerdent, Llama 4 Scout guideundated pageread 9 Oct
- OtherWikipedia, Llama (language model)read 9 Oct
MethodFourteen sources dated 5 April 2025 to 13 July 2026 were read on 9 October 2026 and synthesised with AI, and no model was run and nothing was tested by us.
Dates without a year are 2026.
The review as text
Llama 4 Scout, the Llama-4-Scout-17B-16E-Instruct checkpoint with 17 billion active and 109 billion total parameters, was released by Meta on 5 April 2025. It is the newest Llama-branded model in what we read: two reference pages say Meta moved its main assistant to the closed Muse Spark in April 2026, and neither lists a later Llama. The question here is what developers say they use Scout for beyond its headline.
We read 14 sources on 9 October 2026, dated 5 April 2025 to 13 July 2026. Four are official Meta, Hugging Face and licence pages. Two are an evaluation site and a technology publication. One is a news report about the sibling Maverick model. Five are independent engineering guides and reference pages, and two are developer-forum sources. There is no buyer tier, and the editorial tier is only two pieces, so every count below is small.
Consensus
Across the sources we read, Scout is described as a reader of documents and images, as cheap to call, and as runnable on one large GPU or a high-memory desktop once quantised. The sources dispute the part Meta leads with: the 10-million-token context window. This is our reading of the sample on 9 October 2026, not a measurement of the model.
Confidence is mixed. The facts about size, date and context length come from Meta's own pages, which are the maker's claims. Independent evidence is thin: two evaluation or editorial sources, one dated April 2025.
Recurring strengths
Reading long documents and whole codebases. Three of the 14 sources name it: Meta's launch post (the maker's claim, listing multi-document summarisation and reasoning over large codebases), and two engineering guides dated 8 April and 13 July 2026, one naming document processing and the other whole-codebase analysis and combining many sources in one pass. Neither evaluation source names it as a use.
Reading charts, documents and images. Three sources name it: Meta's model card lists multimodal understanding, one guide names chart interpretation and visual question answering, and the other names text, image and video workflows. The model card adds that image understanding is limited to English.
Cheap to call as a hosted model. Both evaluation or editorial sources point to price. The evaluation site's page, read on 9 October 2026, lists $0.19 per million input tokens and $0.68 per million output tokens, and ranks Scout at 8 on its Intelligence Index, above the 7 it gives as the median of comparable models. The Decoder reported in April 2025 that Scout costs roughly a tenth of GPT-4o. A price is not a quality score.
Runs on one large GPU or a high-memory desktop once quantised. Four sources say so in different terms: the model card (a single H100 with INT4 weights), one guide (48 to 64 GB of VRAM, or a Mac with 64 GB or more, at roughly 12 to 18 tokens a second on those Macs), the other guide (24 GB of VRAM with aggressive quantisation, around 20 tokens a second), and one developer comment reporting almost 20 tokens a second on a Strix Halo machine. The speeds are the authors' own reports.
For scale, the Hugging Face page for Scout showed 170,153 downloads in the last month and 1,368 likes on 9 October 2026. Downloads say how many people fetched the weights, not how many use them.
Recurring complaints
The long-context window does not hold up. Four of the 14 sources raise it. The Decoder reported in April 2025 that on the Fiction.Live benchmark Scout reached 15.6 percent accuracy on narratives of 120,000 tokens. One guide repeats a 15.6 percent figure at 128,000 tokens and puts reliable use near 32,000 tokens. A Verdent guide says the single-H100 claim does not show the full window can be served at production latency, and notes pretraining used 256,000 tokens. In the launch thread on Hacker News, several commenters questioned recall across the whole window. No source counts how many developers hit this, so how common it is stays unknown.
The memory bill. Three sources raise it: a launch-thread commenter putting the minimum near 54.5 GB of VRAM, one guide giving 48 to 64 GB for a quantised setup, and the other guide saying the full 10-million-token window needs eight H100 cards and a budget above $60,000.
Weak for coding. Two sources say so: one guide, which gives Scout 32.8 percent on LiveCodeBench against 33.3 percent for Llama 3.3 70B, and one developer comment saying it fixes bugs by commenting code out and adding TODOs. That comment is one developer's opinion.
Where reviewers split
Whether the long window is a strength. Meta's launch post and one July 2026 guide present it as the reason to pick Scout. The Decoder and the other guide report it failing well before 10 million tokens. We read this as a split between what the window allows and what the benchmark found, and neither side measured the other's case.
General ranking against coding. The evaluation site's index places Scout above its peer median, while one guide calls it unsuitable for coding on one benchmark. These are different tests of different things and both can be true.
A different model, cited once. TechCrunch reported on 11 April 2025 that the public Maverick ranked 32nd on LM Arena, below several older rivals, after Meta had used a tuned experimental version in its launch chart. That is Maverick, not Scout, and is not counted anywhere above.
Individual opinions. A handful of developer comments describe a concrete use, such as local refactoring on a Mac Studio or long-context tests on a private dataset. Each is one developer's account.
Who it suits
Developers with a stack of documents, charts or screenshots to read in bulk, who can either pay a hosted provider a low per-token price or run a quantised copy on a machine with 48 GB or more of memory. Developers who want a long input but will check the answer at lengths past about 32,000 tokens, since the sources we read disagree on how far recall holds.
Who should pass
Anyone who needs dependable recall across hundreds of thousands of tokens, since the sources we read do not support that. Anyone whose main job is writing code, since the two sources that mention coding both fault it. Anyone who wants a model that is still being developed, since the Llama line stopped at this release in what we read. The licence is Meta's Llama 4 Community License Agreement, effective 5 April 2025. This review does not say whether it allows your use, and the licence file is the place to read it.
One verdict a week: the most useful Master Review we finished, the complaint that kept appearing, and who should skip it. Get the weekly verdict
Sources
- Meta, Llama 4 launch post, official, 5 April 2025
- Meta, Llama 4 model card, official, read 9 October 2026
- Hugging Face, Llama-4-Scout-17B-16E-Instruct, official, read 9 October 2026
- Llama 4 Community License Agreement, official, effective 5 April 2025
- Artificial Analysis, Llama 4 Scout, evaluation, read 9 October 2026
- The Decoder, Llama 4 on standard tests and long context, editorial, 7 April 2025
- TechCrunch, vanilla Maverick ranks below rivals, news, 11 April 2025
- Zen van Riel, Llama 4 Scout practical guide, independent guide, 8 April 2026
- DEV Community, Llama 4 Scout locally, independent guide, 13 July 2026
- Verdent, Llama 4 Scout guide, independent guide, undated, read 9 October 2026
- Wikipedia, Llama (language model), reference, read 9 October 2026
- Mungomash, Llama versions, reference, updated 29 September 2026
- Hacker News, The Llama 4 herd, forum, 5 April 2025, partial view of an ordinary page
- Hacker News search, comments on Llama 4 Scout, forum, one API call, read 9 October 2026
Fourteen sources dated 5 April 2025 to 13 July 2026 were read on 9 October 2026 and synthesised with AI, and no model was run and nothing was tested by us.
One verdict a week.
Every week, the most useful Master Review we finished: what the internet agrees on, the complaint that kept appearing, and who should skip it.