Skip to content
All master reviews

Review Machine · Master Review · AI models

Llama 4 Scout

What developers use Meta's newest Llama for

Across 14 sources read on 9 October 2026, Llama 4 Scout is described as a cheap reader of documents and images, while its 10-million-token context window is the part the sources dispute.

Checkpoint Llama-4-Scout-17B-16E-InstructSize 17B active, 109B totalReleased 5 April 2025Licence Llama 4 Community LicenseRead 9 October 2026Skip to the verdictRead as text

Verdict age

ProvisionalTwo evaluations · developers thin

Official pages are from April 2025 and the guides from April to July 2026. A later Llama release, or a new independent long-context evaluation, would change the reading.

  • Early · editors only
  • Provisional · editors in, owners thin
  • Settled · owner reviews read over months

From 14 sources to one verdict

Seven levels, in the same order in every Master Review. Each line is the takeaway; open a level for the evidence under it, or stop when you have enough.

What we read

14 sources, 5 Apr 2025 – 9 Oct 2026. Editorial, other and forum tiers thin. Buyer tier empty.

By tier, heaviest first

T1Official4

Meta's launch post and model card, the Hugging Face page and the licence file.

T2Editorial2Thin

An evaluation page read 9 Oct 2026 and one April 2025 article. Two pieces only.

T2News1

11 Apr 2025

T3Buyers0None readable

No buyer tier exists for a downloadable model.

Forum2Thin

One launch thread as a partial page view and one API call returning nine comments.

Other5Thin

8 Apr – 9 Oct 2026

Date window

5 Apr 2025 to 9 Oct 2026

Editorial reviews 7 Apr 2025 – 9 Oct 2026 · News 11 Apr 2025Launch pages from 5 Apr 2025; guides from 8 Apr 2026 to 13 Jul 2026; pages and figures read 9 Oct 2026.

Couldn’t read, so not used

  • BankInfoSecurity articleHTTP 403
  • GovInfoSecurity articleHTTP 403

Where a page refused a plain fetch, the source was dropped with every count that rested only on it.

Read, then rejected

  • inkeybit Scout guide · gives a month, not a day, so it cannot be dated
  • Maverick spec and price pages · a different model from Scout

What they measured

A claimed 10M-token window, 15.6% accuracy at 120,000 tokens in one benchmark, $0.19 per million input tokens.

Context window claimed by Meta

10Mtokensthe maker's claim

Meta · 5 Apr 2025

Long-narrative benchmark at 120,000 tokens

15.6%Fiction.Live accuracy as reported in April 2025One benchmark

The Decoder · 7 Apr 2025

Hosted price per million tokens

$0.19in

Artificial Analysis · read 9 Oct

$0.68out

Artificial Analysis · read 9 Oct

Intelligence Index

8peer median given as 7Ranking, not use

Artificial Analysis · read 9 Oct

Hugging Face, last month

170,153downloadsNot use

Hugging Face · read 9 Oct

1,368likes

Hugging Face · read 9 Oct

Where they agree

Four uses recur, and only the price is backed by both evaluation sources.

One tick per editorial review, in order of publication:TD The DecoderAA Artificial Analysis

Reading long documents and whole codebases

0/2

Meta's own claim and two independent guides; neither evaluation source names it. T1Official Other

Reading charts, documents and images

0/2

T1Official Other

Cheap to call as a hosted model

2/2

T2Editorial

Runs on one large GPU or a high-memory desktop once quantised

0/2

Speeds are the authors' own reports. T1Official Forum Other

What keeps coming up against it

The long-context window is the complaint that recurs most, in four sources.

The long-context window does not hold up

Documented · size unknown
Documented byT1OfficialT2EditorialOtherForumHow commonUnknown: One benchmark measures it; no source counts how many developers hit it.
T1Official

Meta states a 10M-token window and presents whole-codebase work as a use. This is the maker's claim.

Meta · 5 Apr 2025

T2Editorial

1/2The Decoder reported 15.6 percent on Fiction.Live narratives of 120,000 tokens.

Other

Two guides doubt the window in practice: one puts reliable use near 32,000 tokens, the other says the single-GPU claim does not show full-window serving.

Zen van Riel · 8 AprVerdent · read 9 Oct

Forum

Commenters on the launch thread questioned recall across the whole window.

Hacker News · 5 Apr 2025

Long context fails well before 10M tokens

1/2

T2Editorial Forum Other

The memory bill for local use

0/2

Forum Other

Weak for coding

0/2

One guide's benchmark figure and one developer's comment. Forum Other

Where they split

Sources split on whether the long window is a strength.

Is the 10M-token window a reason to pick Scout?

Meta, DEV Community

Presented as the reason to pick it, for codebases and many documents in one pass.

T2The Decoder, Zen van Riel

Against that: Reported failing well before 10 million tokens, with a benchmark figure.

Our readThe window is what the model accepts. Whether answers stay reliable is what the benchmark questions.

General ranking against coding

T2Artificial Analysis

Ranks Scout above its peer median on its index.

Zen van Riel

Against that: Calls it unsuitable for coding on one benchmark.

Both trueDifferent tests of different things.

A different model: Maverick

T1Meta launch chart

A tuned experimental Maverick is shown on a public chat leaderboard.

T2TechCrunch, 11 Apr 2025

Against that: The public Maverick ranked 32nd, below several older rivals. Not Scout, and not counted.

Who it's for, who should pass

Suits bulk reading of documents and images; not long-recall or coding work.

It suits you if

  • You have documents, charts or screenshots to read in bulk and pay per token.Price and image reading recur in the sources.
  • You have a machine with 48 GB or more of memory and want open weights.Two guides and the model card give these requirements.
  • You want a long input but will check answers past about 32,000 tokens.The sources disagree on how far recall holds.

Pass if

  • You need dependable recall across hundreds of thousands of tokens.The Decoder and two guides report it failing.
  • Your main work is writing code.A guide and a developer comment both fault it.
  • You want a model that is still being developed.Two reference pages say Meta moved to Muse Spark in April 2026.

The verdict

Used for reading and price, disputed on its headline feature.
VerdictProvisionalTwo evaluations · developers thin

By the sources we read, Scout is a cheap reader of documents and images whose 10-million-token window is the part they dispute, from a small sample.

Confidence, by tier

Officialstrong

Size, date and claims come straight from Meta's pages.

Editorialthin

Two pieces, one from launch week.

Otherthin

Three guides by individual authors.

Forumthin

A partial thread view and nine comments.

Buyersnone

No buyer tier exists.

Editorial evidence
up to 18 months old
Newest report
11 Apr 2025
Read
9 Oct · month 18

Share line

Ten million tokens of context, and a benchmark found trouble at 120,000.
Review MachineLlama 4 Scout · Master Review

Rests onThe Decoder, April 2025: Scout reached 15.6 percent on Fiction.Live narratives of 120,000 tokens, against Meta's claimed 10M-token window.

Sources

14 sources, heaviest tier first. Every figure above comes from one of these.

  1. T1OfficialMeta, Llama 4 launch post5 Apr 2025
  2. T1OfficialLlama 4 Community License Agreementupdated 5 Apr 2025
  3. T1OfficialMeta, Llama 4 model cardread 9 Oct
  4. T1OfficialHugging Face, Llama-4-Scout-17B-16E-Instructone page load, headline figures onlyread 9 Oct
  5. T2EditorialThe Decoder, Llama 4 on standard tests and long context7 Apr 2025
  6. T2EditorialArtificial Analysis, Llama 4 Scoutevaluation page, one loadread 9 Oct
  7. T2NewsTechCrunch, vanilla Maverick ranks below rivals11 Apr 2025
  8. ForumHacker News, The Llama 4 herdpartial view of an ordinary page5 Apr 2025
  9. ForumHacker News search, comments on Llama 4 Scoutone API call, nine comments shownread 9 Oct
  10. OtherZen van Riel, Llama 4 Scout practical guide8 Apr
  11. OtherDEV Community, Llama 4 Scout locally13 Jul
  12. OtherMungomash, Llama versionsupdated 29 Sep
  13. OtherVerdent, Llama 4 Scout guideundated pageread 9 Oct
  14. OtherWikipedia, Llama (language model)read 9 Oct

MethodFourteen sources dated 5 April 2025 to 13 July 2026 were read on 9 October 2026 and synthesised with AI, and no model was run and nothing was tested by us.

Dates without a year are 2026.

The review as text

The same review as one piece of writing · 7 min read

Llama 4 Scout, the Llama-4-Scout-17B-16E-Instruct checkpoint with 17 billion active and 109 billion total parameters, was released by Meta on 5 April 2025. It is the newest Llama-branded model in what we read: two reference pages say Meta moved its main assistant to the closed Muse Spark in April 2026, and neither lists a later Llama. The question here is what developers say they use Scout for beyond its headline.

We read 14 sources on 9 October 2026, dated 5 April 2025 to 13 July 2026. Four are official Meta, Hugging Face and licence pages. Two are an evaluation site and a technology publication. One is a news report about the sibling Maverick model. Five are independent engineering guides and reference pages, and two are developer-forum sources. There is no buyer tier, and the editorial tier is only two pieces, so every count below is small.

Consensus

Across the sources we read, Scout is described as a reader of documents and images, as cheap to call, and as runnable on one large GPU or a high-memory desktop once quantised. The sources dispute the part Meta leads with: the 10-million-token context window. This is our reading of the sample on 9 October 2026, not a measurement of the model.

Confidence is mixed. The facts about size, date and context length come from Meta's own pages, which are the maker's claims. Independent evidence is thin: two evaluation or editorial sources, one dated April 2025.

Recurring strengths

Reading long documents and whole codebases. Three of the 14 sources name it: Meta's launch post (the maker's claim, listing multi-document summarisation and reasoning over large codebases), and two engineering guides dated 8 April and 13 July 2026, one naming document processing and the other whole-codebase analysis and combining many sources in one pass. Neither evaluation source names it as a use.

Reading charts, documents and images. Three sources name it: Meta's model card lists multimodal understanding, one guide names chart interpretation and visual question answering, and the other names text, image and video workflows. The model card adds that image understanding is limited to English.

Cheap to call as a hosted model. Both evaluation or editorial sources point to price. The evaluation site's page, read on 9 October 2026, lists $0.19 per million input tokens and $0.68 per million output tokens, and ranks Scout at 8 on its Intelligence Index, above the 7 it gives as the median of comparable models. The Decoder reported in April 2025 that Scout costs roughly a tenth of GPT-4o. A price is not a quality score.

Runs on one large GPU or a high-memory desktop once quantised. Four sources say so in different terms: the model card (a single H100 with INT4 weights), one guide (48 to 64 GB of VRAM, or a Mac with 64 GB or more, at roughly 12 to 18 tokens a second on those Macs), the other guide (24 GB of VRAM with aggressive quantisation, around 20 tokens a second), and one developer comment reporting almost 20 tokens a second on a Strix Halo machine. The speeds are the authors' own reports.

For scale, the Hugging Face page for Scout showed 170,153 downloads in the last month and 1,368 likes on 9 October 2026. Downloads say how many people fetched the weights, not how many use them.

Recurring complaints

The long-context window does not hold up. Four of the 14 sources raise it. The Decoder reported in April 2025 that on the Fiction.Live benchmark Scout reached 15.6 percent accuracy on narratives of 120,000 tokens. One guide repeats a 15.6 percent figure at 128,000 tokens and puts reliable use near 32,000 tokens. A Verdent guide says the single-H100 claim does not show the full window can be served at production latency, and notes pretraining used 256,000 tokens. In the launch thread on Hacker News, several commenters questioned recall across the whole window. No source counts how many developers hit this, so how common it is stays unknown.

The memory bill. Three sources raise it: a launch-thread commenter putting the minimum near 54.5 GB of VRAM, one guide giving 48 to 64 GB for a quantised setup, and the other guide saying the full 10-million-token window needs eight H100 cards and a budget above $60,000.

Weak for coding. Two sources say so: one guide, which gives Scout 32.8 percent on LiveCodeBench against 33.3 percent for Llama 3.3 70B, and one developer comment saying it fixes bugs by commenting code out and adding TODOs. That comment is one developer's opinion.

Where reviewers split

Whether the long window is a strength. Meta's launch post and one July 2026 guide present it as the reason to pick Scout. The Decoder and the other guide report it failing well before 10 million tokens. We read this as a split between what the window allows and what the benchmark found, and neither side measured the other's case.

General ranking against coding. The evaluation site's index places Scout above its peer median, while one guide calls it unsuitable for coding on one benchmark. These are different tests of different things and both can be true.

A different model, cited once. TechCrunch reported on 11 April 2025 that the public Maverick ranked 32nd on LM Arena, below several older rivals, after Meta had used a tuned experimental version in its launch chart. That is Maverick, not Scout, and is not counted anywhere above.

Individual opinions. A handful of developer comments describe a concrete use, such as local refactoring on a Mac Studio or long-context tests on a private dataset. Each is one developer's account.

Who it suits

Developers with a stack of documents, charts or screenshots to read in bulk, who can either pay a hosted provider a low per-token price or run a quantised copy on a machine with 48 GB or more of memory. Developers who want a long input but will check the answer at lengths past about 32,000 tokens, since the sources we read disagree on how far recall holds.

Who should pass

Anyone who needs dependable recall across hundreds of thousands of tokens, since the sources we read do not support that. Anyone whose main job is writing code, since the two sources that mention coding both fault it. Anyone who wants a model that is still being developed, since the Llama line stopped at this release in what we read. The licence is Meta's Llama 4 Community License Agreement, effective 5 April 2025. This review does not say whether it allows your use, and the licence file is the place to read it.

One verdict a week: the most useful Master Review we finished, the complaint that kept appearing, and who should skip it. Get the weekly verdict

Sources

Fourteen sources dated 5 April 2025 to 13 July 2026 were read on 9 October 2026 and synthesised with AI, and no model was run and nothing was tested by us.

One verdict a week.

Every week, the most useful Master Review we finished: what the internet agrees on, the complaint that kept appearing, and who should skip it.

By subscribing you agree to receive one weekly email from Review Machine. You can unsubscribe at any time.