Skip to content
All master reviews

Review Machine · Master Review · ai-models

DeepSeek R1 and the R1-0528 checkpoint

Across fifteen sources read on 24 September 2026 and dated 20 January 2025 to 24 April 2026, DeepSeek R1 was received in January 2025 as an open-weights shock, checked in May by one independent index that moved 60 to 68 while thinking tokens rose 40%, and left behind in 2026 by a maker that retired the API alias serving it on 24 July 2026 while the MIT-licensed weights stayed published.

Released 20 January 2025, MIT licenceCurrent checkpoint R1-0528, 28 May 2025Size 671B total, 37B activeAPI alias deepseek-reasoner, retired 24 July 2026Read 24 September 2026Skip to the verdictRead as text

Verdict age

SettledTwenty months of coverage · one independent test

The January 2025 coverage is twenty months old and the last independent test is sixteen months old. What could change is a rerun: no source in this reading has yet tested the published weights against the maker's own table.

  • Early · editors only
  • Provisional · editors in, owners thin
  • Settled · owner reviews read over months

From 15 sources to one verdict

Seven levels, in the same order in every Master Review. Each line is the takeaway; open a level for the evidence under it, or stop when you have enough.

What we read

15 sources, 20 Jan 2025 – 24 Sep 2026. Editorial, buyer and forum tiers thin. Video tier empty.

By tier, heaviest first

T1Official5

Five DeepSeek pages: the R1 release note, both model cards, the paper and the V4 preview note.

T2Editorial2Thin

Thin. Two independent write-ups, and only one of them ran tests of its own.

T2News4

Four reports, from January 2025 to September 2025.

T3Buyers1Thin

One GitHub REST API call, reported as that API's figures on the day read.

Forum1Thin

One Hacker News API search, a count of stories rather than a reading of them.

Other2

Two security firms' own reports of their own testing.

Date window

20 Jan 2025 to 24 Sep 2026

Editorial reviews 22 Jan – 29 May 2025 · News 27 Jan – 19 Sep 202520 January 2025 to 24 April 2026, all read on 24 September 2026.

Couldn’t read, so not used

  • LiveCodeBench leaderboardtable loaded from script and stayed empty on the read

Nothing here rests only on the leaderboard. The coding figure in this piece is the SCMP's report of it, not our reading of the table.

Read, then rejected

  • deepseekai.guide R1 review · a guide site's review with no named author, dated only by a caution line

What they measured

The maker's own table moved furthest in May 2025, and the one independent index we could read moved with it while adding a 40% rise in thinking tokens.

AIME 2025, the maker's card

Average thinking tokens per AIME question

23KtokensR1-0528, against 12K for R1, as the maker's card states it

Hugging Face · read 24 Sep

Artificial Analysis Intelligence Index

68pointsR1-0528, up from 60 for the January checkpoint

Artificial Analysis · 29 May 2025

Hacker News stories mentioning DeepSeek-R1

374storiesOne Algolia API search, run 24 September 2026

Hacker News API · read 24 Sep

deepseek-ai/DeepSeek-R1 on GitHub

91,975starsOne GitHub API call, 24 September 2026; last push 27 June 2025

GitHub REST API · read 24 Sep

Where they agree

The release, the May refresh and the 2026 retirement are three separate events, and the reading gives each of them its own date.

One tick per editorial review, in order of publication:TB The BatchAA Artificial Analysis

The licence, not the benchmark score, was what the January coverage called new

1/2

All three name the MIT-licensed weights as the substance of the release, from 20 to 27 January 2025. T1Official T2Editorial T2News

The May refresh moved the numbers without touching the architecture

1/2

A post-training update on the same 671B model, reported by the maker, the SCMP and an independent index on 29 May 2025. T1Official T2Editorial T2News

Thinking got longer as the answers got better

1/2

Average AIME thinking rose from 12K to 23K tokens, and the independent index cost 99 million tokens against 71 million. T1Official T2Editorial

By 2026 the maker had moved on and the weights had not

0/2

The legacy alias retired on 24 July 2026 while the repository's last push was 27 June 2025. T1Official T3Buyers

What keeps coming up against it

The complaint that recurs is safety and exposure, documented by two firms inside ten days of the launch and never sized by anyone.

Safety and exposure findings in the first ten days

Documented by two firms · no source counted how common it is
Documented byCiscoT2SecurityWeek · 4 Feb 2025WizHow commonUnknown: Three sources document findings dated January and February 2025. None counts how often the weakness is reached, so this is not a rate.
Cisco

Cisco's security researchers, with the University of Pennsylvania, ran an automated jailbreak on 50 prompts from HarmBench and reported a 100% attack success rate against R1.

  • SecurityWeek The same run put OpenAI's o1-preview at 26% and the other models between 36% and 96%.
  • Cisco The report links the outcome to training methods that were optimised for cost.

Cisco · 31 Jan 2025

T2SecurityWeek · 4 Feb 2025

SecurityWeek reported the comparison across six models and recorded R1 as the only one no jailbreak attempt failed against.

SecurityWeek · 4 Feb 2025

Wiz

Wiz Research reported a publicly accessible and unauthenticated database belonging to DeepSeek, holding over a million log entries including chat history and API keys, and says DeepSeek secured it promptly.

Wiz · 29 Jan 2025

The maker's headline claims were never independently rerun

0/2

Ars Technica reported the cost figure as an allegation in January 2025, and Reuters reported in September 2025 that US officials had questioned the compute picture. T2News

Where they split

The two splits that matter are what the maker's own table is worth, and which way the May refresh cut.

What the January benchmark table is worth

T1DeepSeek · paper, 22 Jan 2025

The paper reports R1 as comparable to o1 across math, code and reasoning tasks, and the model card reports the distilled 32B model outperforming o1-mini.

T2Ars Technica · 27 Jan 2025

Against that: Ars Technica took the cost figure as an allegation and said it would evaluate R1 more formally later. No independent rerun of the January table is in this sample.

Our readThe January table is the maker's own, and it became checkable only in May, when an independent index and a leaderboard published numbers of their own.

Which way the May refresh cut

T1DeepSeek · model card

The card presents a minor version upgrade with a reduced hallucination rate and gains on AIME 2025, GPQA-Diamond and SWE Verified.

T2Artificial Analysis · 29 May 2025

Against that: Its index moved 60 to 68, and it counted 99 million tokens of thinking for the refresh against 71 million for the January checkpoint.

Both trueThe card's own SimpleQA row also falls, from 30.1 to 27.8, inside the same table that reports the reduced hallucination rate.

Who it's for, who should pass

It still suits anyone building on open weights, and it no longer suits anyone who needs a current endpoint or a guarded one.

It suits you if

  • You want reasoning weights you can download, fine-tune and distil.DeepSeek's release page lists the code and models under the MIT licence with distillation and commercial use allowed, and the repository is still public under it on 24 September 2026.
  • You want to see the working rather than the answer alone.The Batch noted on 22 January 2025 that R1 shows the chain of thought it works through.
  • You want a small model that inherited the reasoning.Artificial Analysis scored the distilled R1-0528-Qwen3-8B at 52 on its index, and the maker's card reports it matching a 235B thinking model on AIME 2024.
  • You are still calling deepseek-reasoner in code you maintain.The alias retired on 24 July 2026 and now routes to a newer model.

Pass if

  • You need the current frontier rather than a 2025 milestone.The repository's last push was 27 June 2025 and the alias that served the model retired on 24 July 2026.
  • You need guardrails out of the box.The two security findings above are dated 29 and 31 January 2025, and no source in this reading revisited them.
  • You need the maker's cost and benchmark claims verified before you commit.One independent evaluation and one leaderboard report carry the whole of the verification in this sample.

The verdict

A milestone whose weights outlived its endpoint.
VerdictSettledTwenty months of coverage · one independent test

R1 is still worth reading as the moment open weights were taken seriously, and it is no longer the model to build a product on, because its maker retired the endpoint and nobody has rerun its January numbers.

Confidence, by tier

Officialstrong

Five DeepSeek pages, including both model cards and the release note that set the licence.

Editorialthin

One independent index that ran its own tests, and one write-up that summarised the launch.

Newsstrong

Four reports from February to September 2025, including the peer-reviewed cost coverage.

Buyersthin

One GitHub API call. Stars are attention, not a satisfaction figure.

Forumthin

One API count of stories. Nothing was counted across the threads themselves.

Otherstrong

Two security firms reporting their own testing, each on its own date.

Videonone

No video source was read.

Editorial evidence
15–20 months old
Newest report
19 Sep 2025
Read
24 Sep · month 20

Share line

The weights are still free. The endpoint that served them is not.
Review MachineDeepSeek R1 and the R1-0528 checkpoint · Master Review

Rests onDeepSeek's V4 preview note, 24 April 2026, sets the retirement of the deepseek-reasoner alias at 24 July 2026, while both model cards stay published under MIT.

Sources

15 sources, heaviest tier first. Every figure above comes from one of these.

  1. T1OfficialDeepSeek-R1 Release20 Jan 2025
  2. T1OfficialDeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning22 Jan 2025
  3. T1OfficialDeepSeek-V4 Preview Release24 Apr
  4. T1Officialdeepseek-ai/DeepSeek-R1 model cardread 24 Sep
  5. T1Officialdeepseek-ai/DeepSeek-R1-0528 model cardread 24 Sep
  6. T2EditorialDeepSeek-R1, an affordable rival to OpenAI's o122 Jan 2025
  7. T2EditorialDeepSeek's R1 leaps over xAI, Meta and Anthropic29 May 2025
  8. T2NewsDeepSeek panic triggers tech stock sell-off as Chinese AI tops App Store27 Jan 2025
  9. T2NewsDeepSeek Compared to ChatGPT, Gemini in AI Jailbreak Test4 Feb 2025
  10. T2NewsDeepSeek quietly updates R1 AI model29 May 2025
  11. T2NewsChina's DeepSeek shook the tech world19 Sep 2025
  12. T3Buyersdeepseek-ai/DeepSeek-R1 repository figuresone API call: stars, forks, open issues, licence, last pushread 24 Sep
  13. ForumStory search for DeepSeek-R1read 24 Sep
  14. OtherWiz Research Uncovers Exposed DeepSeek Database29 Jan 2025
  15. OtherEvaluating Security Risk in DeepSeek31 Jan 2025

MethodMethod: 15 sources dated 20 January 2025 to 24 April 2026, read on 24 September 2026 and synthesised with AI. Nothing was run, prompted, installed or tested.

Dates without a year are 2026.

The review as text

The same review as one piece of writing · 7 min read

DeepSeek R1 is DeepSeek's first-generation reasoning model, published as open weights on 20 January 2025 under the MIT licence, with the R1-0528 checkpoint of 28 May 2025 as its current published version. We read 15 sources on 24 September 2026, dated 20 January 2025 to 24 April 2026: five pages published by DeepSeek itself, two independent write-ups, four news reports, two security firms' reports, and one public API read each of Hacker News and GitHub.

The editorial tier is thin for a subject like this one. No publication we read ran its own laboratory test of R1, and one independent index carries most of the verification. The buyer tier is the model's own audience: GitHub's API reported 91,975 stars, 11,679 forks and 37 open issues for deepseek-ai/DeepSeek-R1 on 24 September 2026, with its last push on 27 June 2025. Stars measure attention, not quality.

Consensus

The reading moved in three stages, and the sample supports each plainly.

In January 2025, three of the sources we read treated the release itself as the event. DeepSeek's own release page put the code and the models under the MIT licence, with distillation and commercial use allowed. The Batch, on 22 January 2025, read that licence as the part that could push the state of the art forward at every size. Ars Technica, on 27 January 2025, counted the free, openly licensed weights as one of three things that shocked the experts it spoke to. One search of the public Hacker News API returned 374 stories mentioning the model on 24 September 2026, which is a figure for attention rather than approval.

In May 2025, three sources recorded a second stage. DeepSeek republished the model as R1-0528 with no change to its architecture, still 671 billion parameters with 37 billion active. Its own card reports AIME 2025 rising from 70.0 to 87.5 and GPQA-Diamond from 71.5 to 81.0. The SCMP, on 29 May 2025, recorded that the checkpoint went up on Hugging Face with no announcement and no documentation, and that LiveCodeBench reported a coding gain leaving it behind only three OpenAI models. Artificial Analysis, the same day, moved the model from 60 to 68 on its own index, tying it with Gemini 2.5 Pro.

In 2026 the two halves separated. DeepSeek's release page for its V4 preview, dated 24 April 2026, set the retirement of the legacy API aliases at 24 July 2026 and noted that calls to them already route to a newer model. The repository did not follow: its last push was 27 June 2025, and it is still public under MIT.

Confidence is strong on what DeepSeek published and what its API does, and thin on the model's behaviour, because one independent index and one leaderboard report carry all of it.

Recurring strengths

Three strengths recur, and two of them are about the release rather than the answers.

The licence. Three sources name it as the release's substance: DeepSeek's release page, The Batch on 22 January 2025 and Ars Technica on 27 January 2025. Because it permits distillation, it was the route by which R1 reached sizes nobody could serve at full size.

The visible working. The Batch, on 22 January 2025, noted that R1 shows the chain of thought it works through, where the reasoning models it was compared with did not. That made the reasoning readable as text, which is what much of the January coverage was examining.

Thinking got longer as the answers got better. DeepSeek's card for R1-0528 reports average thinking length on AIME rising from 12K tokens to 23K. Artificial Analysis counted 99 million tokens to complete its index, against 71 million for the January checkpoint, a rise of 40%.

Recurring complaints

The recurring complaint is safety and exposure, and it arrived inside the first ten days. Cisco's security team, with the University of Pennsylvania, ran an automated jailbreak on 50 prompts from HarmBench and reported a 100% attack success rate against R1, against 26% for OpenAI's o1-preview in the same run. SecurityWeek reported that comparison on 4 February 2025. Wiz Research, on 29 January 2025, reported a publicly accessible and unauthenticated database belonging to DeepSeek, one exposure holding over a million log entries including chat history and API keys, and says DeepSeek secured it promptly. Neither finding states how often the weakness is reached in practice, and neither is a rate.

A second complaint runs the whole length of the window: the maker's headline claims were never independently rerun. Ars Technica reported the training-cost figure as an allegation on 27 January 2025. Reuters reported on 19 September 2025 that the cost and compute picture had been questioned by US officials, and that the peer-reviewed article in Nature carried the first cost estimate DeepSeek had given, 294,000 dollars for R1 on 512 H800 chips.

Where reviewers split

Two splits are worth holding apart.

On what the January benchmark table is worth. DeepSeek's own paper and card report R1 as comparable to o1 across math, code and reasoning tasks. Ars Technica, on 27 January 2025, took the cost figure as an allegation and said it would evaluate R1 more formally later. Our read: the January table is the maker's own, and it became checkable only in May, when two independent bodies published numbers of their own.

On which way the May refresh cut. DeepSeek's card presents it as a minor upgrade with a reduced hallucination rate. Artificial Analysis scored it 68 on its own index and counted the 40% rise in thinking tokens. Both are true, and the card's own SimpleQA row falls, from 30.1 to 27.8, inside the table that reports the reduced hallucination rate.

Who it suits

You want reasoning weights you can download, fine-tune and distil without asking anyone. DeepSeek's release page lists the code and the models under the MIT licence with distillation and commercial use allowed, and the repository is still public under the same licence on 24 September 2026.

You want to see the working. The Batch noted on 22 January 2025 that R1 shows its chain of thought rather than hiding it.

You want a small model that inherited the reasoning. Artificial Analysis scored the distilled R1-0528-Qwen3-8B at 52 on its index, and DeepSeek's card reports that model matching a 235B thinking model on AIME 2024.

You are still calling deepseek-reasoner in code you maintain. That alias retired on 24 July 2026 and now routes to a newer model.

Who should pass

You need the current frontier rather than a 2025 milestone. The repository's last push was 27 June 2025, and the alias that served the model was retired on 24 July 2026.

You need guardrails out of the box. The two security findings above are dated 29 January and 31 January 2025, and no source in this reading revisited them.

You need the maker's cost and benchmark claims verified before you commit. One independent evaluation and one leaderboard report carry the whole of the verification.

One verdict a week: the most useful Master Review we finished, the complaint that kept appearing, and who should skip it, at the weekly verdict.

Sources

Method: 15 sources dated 20 January 2025 to 24 April 2026, read on 24 September 2026 and synthesised with AI. Nothing was run, prompted, installed or tested.

One verdict a week.

Every week, the most useful Master Review we finished: what the internet agrees on, the complaint that kept appearing, and who should skip it.

By subscribing you agree to receive one weekly email from Review Machine. You can unsubscribe at any time.