Skip to content
All master reviews

Review Machine · Master Review · AI libraries

LiteLLM

Every README is a promise. The issues tab reports back.

Across the README, the licence, the spend-tracking docs, one repository API call, one issue-search call per theme, five open issues, two developer threads and two editorial assessments read on 18 September 2026, LiteLLM v1.101.0 is the broadest self-hosted gateway with its value in keys, budgets and logs, and its recurring costs are provider behaviour that does not fit the OpenAI shape, spend accounting that takes work, and a proxy that is not light.

Release read v1.101.0, 15 September 2026Licence MIT outside the enterprise directoryRead 18 September 2026Skip to the verdictRead as text

Verdict age

ProvisionalEditors in · developers thin

The editorial tier is 5 months and 16 months old, and the incident is 6 months old. The issues, the repository figures and the release list were read on 18 September 2026, three days after v1.101.0, so the complaint counts move with every release.

  • Early · editors only
  • Provisional · editors in, owners thin
  • Settled · owner reviews read over months

From 21 sources to one verdict

Seven levels, in the same order in every Master Review. Each line is the takeaway; open a level for the evidence under it, or stop when you have enough.

What we read

21 sources, 15 May 2025 – 18 Sep 2026. Editorial, news, forum and other tiers thin.

By tier, heaviest first

T1Official5

The README, the licence file, the spend-tracking page, the release list and the maker's own benchmark post, read 18 Sep 2026.

T2Editorial2Thin

Two assessments, 15 May 2025 and 15 Apr 2026; the older one predates the Rust gateway and the March 2026 incident.

T2News1Thin

One report of the 24 March 2026 PyPI compromise.

T3Buyers2

One repository API call and one issue-search API call per theme, figures as returned on 18 Sep 2026.

Forum8Thin

Five open issues and two developer threads read as ordinary pages; 563 comments between the two threads.

Other3Thin

Three comparison and benchmark write-ups, read 18 Sep 2026.

Date window

15 May 2025 to 18 Sep 2026

Editorial reviews 15 May 2025 – 15 Apr 2026 · News 24 Mar 2026Pages with no date of their own are dated 18 Sep 2026, the day they were read.

Couldn’t read, so not used

  • 5 Real Issues With LiteLLM (dev.to)HTTP 404 when fetched

The one page that could not be opened is left out and no count rests on it. No issue list was paged, no site filter was applied, and each open-issue total comes from one search API call.

Read, then rejected

  • CompsMag, LiteLLM Review 2026 · publishes its own test results, which we cannot check; nothing was counted from it
  • Future AGI, best LiteLLM alternatives · vendor comparison, little on the themes counted
  • Endor Labs, TeamPCP research · cited through BleepingComputer, not opened directly

What they measured

59,017 stars, 11,538 forks and 5,201 open issues on 18 September 2026; 102, 71 and 18 of those issues name streaming, cost and memory in their titles.

Repository, 18 Sep 2026

59,017stars

GitHub · read 18 Sep

11,538forks

GitHub · read 18 Sep

5,201open issues

GitHub · read 18 Sep

Open issues naming a word in the title

102issuesstreaming

GitHub · read 18 Sep

71issuescost

GitHub · read 18 Sep

18issuesmemory

GitHub · read 18 Sep

One search API call per word, 18 September 2026. An issue count is not a defect rate.

Added latency, the maker's own test

0.7msRust gateway, p99Maker's own benchmark

LiteLLM docs · 22 Jul

257.7mslegacy Python proxy, p99Same test, mock upstream

LiteLLM docs · 22 Jul

Peak memory, the same test

21.8MBRust gateway

LiteLLM docs · 22 Jul

329.5MBlegacy Python proxy

LiteLLM docs · 22 Jul

Time to first byte under load

35.49msat 5 virtual users

Ravish Tiwari · read 18 Sep

163.02msat 10 virtual users, same write-upIndependent write-up

Ravish Tiwari · read 18 Sep

Versions affected in March 2026

1.82.7published to PyPI, removed

BleepingComputer · 24 Mar

1.82.8published to PyPI, removed

BleepingComputer · 24 Mar

Where they agree

One OpenAI-shaped call over more than a hundred providers, and the value in the keys, the budgets and the logs, in the editorial assessments and the comparisons alike.

One tick per editorial review, in order of publication:IW InfoWorldTH Thoughtworks

One shaped call over many providers

2/2

Both editorial assessments and three comparisons describe the same core. T1Official T2Editorial Other

The value sits in keys, budgets and logs

2/2

Virtual keys scoped per team or app, budgets and rate limits on them, request logs and spend tracking. T1Official T2Editorial Other

It ships often

0/2

v1.101.0 was published on 15 September 2026, and the repository's last push was 17 September 2026. T1Official T3Buyers

The licence split is stated in the open

0/2

MIT outside the enterprise directory; its own licence inside it. T1Official

GitHub's own figures

GitHub's figures for the repository, one API call: 59,017 stars, 11,538 forks and 5,201 open issues on 18 September 2026. T3Buyers

5201open issues59017stars11538forks

GitHub's own counts, from one API call. Stars are attention, not quality.

What keeps coming up against it

Provider behaviour that does not fit the shape, spend accounting that takes work, a proxy that is not light, and two compromised releases.

Provider behaviour that does not fit the OpenAI shape

Well documented · size unknown
Documented byT2EditorialForumT3BuyersT1OfficialHow commonUnknown: no source we read counts how common it is; 102 open issues name streaming in the title, and an issue count is not a rate
T2Editorial

1/2Its teams adopt the gateway as a default, and it reports that teams relying on a provider's own capabilities end up passing provider-specific parameters anyway, reintroducing the coupling the gateway exists to remove; its entry also says drop_params mode discards unsupported parameters silently.

Forum

Open issues describe streaming-specific gaps.

  • GitHub reasoning text dropped when a Responses-shaped answer is bridged to Chat Completions, streaming and non-streaming alike
  • GitHub Gemini grounding metadata intermittently dropped on streaming responses

GitHub · 11 SepGitHub · 16 Sep

T3Buyers

One search API call returns 102 open issues with the word streaming in the title on 18 September 2026. Count, not a rate

GitHub · read 18 Sep

T1Official

The README sells a single OpenAI-shaped call over more than a hundred providers.

BerriAI · read 18 Sep

Cost and spend accounting takes work

0/2

71 open issues name cost in the title on 18 September 2026, and the project's own spend-tracking page documents that a self-declared user id can move spend off the key. T1Official T3Buyers Forum

It is not light, whatever the name says

0/2

18 open issues name memory in the title, and the developer thread of 11 September 2026 is about a fork built by cutting features out, with the comments naming the cut features as the reason people run the gateway. T1Official T3Buyers Forum Other

Two compromised releases on PyPI

0/2

1.82.7 and 1.82.8, 24 March 2026, reported by BleepingComputer from Endor Labs research, with the maintainer's own statement in the same thread. T2News Forum

GitHub's open-issue count

5,201 open issues on 18 September 2026, one API call. The two most-named words in their titles are streaming and cost, at 102 and 71. T3Buyers Counts only

GitHub's own counts, not ours. An issue count is not a defect rate.

Where they split

Three places the evidence pulls two ways: what the benchmarks measure, how far the incident reached, and what the weight is worth.

Performance · what a benchmark measures

T1LiteLLM

Its own post reports 0.7 ms of p99 added latency and 21.8 MB of peak memory for the Rust gateway, against 257.7 ms and 329.5 MB for the legacy Python proxy, on a local mock upstream.

DeepInspect · 9 Sep

Against that: Both readings say a mock upstream deletes the variable that dominates real requests, and that vendor overhead figures are not comparable with each other.

Our readNeither side says what a gateway adds on a reader's own traffic.

How far the incident reached

The maintainer

The official Docker proxy path pins its dependencies and was not affected by the two packages; accounts and keys were rotated and releases paused while the release path was reviewed.

T2BleepingComputer

Against that: Roughly 500,000 exfiltration events, which it reports it could not confirm independently, from research by Endor Labs.

SoThe exposure turns on how the gateway was installed, and none of the sources we read settled it.

Whether the weight is worth paying

Forum

The complaint is the weight and the feature list, and the comments about a stripped-down fork argue that the features being cut are the reason people run it.

Medium benchmark

Against that: It puts LiteLLM ahead on provider, endpoint and administration coverage while it loses on latency.

Both trueThe same breadth is the weight and the reason it stays in the stack.

Who it's for, who should pass

For teams with their own provider contracts who want budgets and logs they control. Pass if speed is the reason, or the install cannot be pinned.

It suits you if

  • You already hold provider contracts and want one OpenAI-shaped endpoint in front of them.Both editorial assessments describe that as the case it is built for.
  • You want per-key budgets, scoped keys and request logs you run yourself.The README and both editorial assessments put the value there, and the spend-tracking page shows the controls.
  • You will pin versions and read release notes.The two compromised releases of 24 March 2026 turned on what an unpinned install could pick up.
  • You need a provider's own capabilities and can configure them explicitly.Thoughtworks reports they do not survive the translation on their own.

Pass if

  • You are choosing it for raw speed on the strength of a benchmark.No test without a product in the comparison turned up, and the independent readings say those benchmarks cannot be compared.
  • You expect a provider's own response fields to arrive unchanged through a streaming bridge.102 open issues name streaming in the title on 18 September 2026.
  • The install cannot be pinned to known-good versions.1.82.7 and 1.82.8 were published to PyPI on 24 March 2026 and reported carrying a credential stealer.
  • You want the smallest possible dependency tree.18 open issues name memory in the title, and the developer thread of 11 September 2026 is a fork built by cutting features out.

This is a reading of published reviews, not medical, financial or legal advice. For a decision about your health, your money or your rights, a qualified professional is the right next step, and not a review.

The verdict

Broad and governed, and heavier and more brittle at the edges than its README suggests.
VerdictProvisionalEditors in · developers thin

A gateway that does put more than a hundred providers behind one OpenAI-shaped call and does put the control in keys, budgets and logs, with three repeated costs: provider behaviour that does not fit the shape, spend accounting that takes work, and a proxy that is not light.

Confidence, by tier

Officialstrong

README, licence, spend-tracking docs, release list and the maker's own benchmark post, read 18 Sep 2026.

Editorialthin

Two assessments, one 5 months old and one 16 months old, and neither is a full review of this release.

Newsthin

One report, of the March 2026 PyPI compromise.

Buyersstrong

GitHub's figures and the open-issue totals, as returned. They measure attention and volume, not quality.

Forumthin

Five issues and two threads, 563 comments between the threads.

Otherthin

Three write-ups, two of them from parties selling a competing gateway.

Editorial evidence
5–16 months old
Newest report
24 Mar 2026
Read
18 Sep · month 41

Share line

One endpoint for 100 providers. The hundred and first is your provider's own quirks.
Review MachineLiteLLM · Master Review

Rests on102 open issues with streaming in the title on 18 September 2026, and the Thoughtworks Technology Radar entry updated 15 April 2026 on provider-specific parameters.

Sources

21 sources, heaviest tier first. Every figure above comes from one of these.

  1. T1OfficialBenchmarking the LiteLLM Rust AI Gateway22 Jul
  2. T1OfficialRelease v1.101.015 Sep
  3. T1OfficialLiteLLM READMEread 18 Sep
  4. T1OfficialLiteLLM licence fileread 18 Sep
  5. T1OfficialSpend Trackingread 18 Sep
  6. T2EditorialLiteLLM: An open-source gateway for unified LLM access15 May 2025
  7. T2EditorialLiteLLM, Technology Radarupdated 15 Apr
  8. T2NewsPopular LiteLLM PyPI package backdoored to steal credentials24 Mar
  9. T3BuyersLiteLLM repository figuresone API call, figures as returnedread 18 Sep
  10. T3BuyersGitHub issue searchone search API call per theme, open-issue totals as returnedread 18 Sep
  11. ForumTell HN: Litellm 1.82.7 and 1.82.8 on PyPI are compromised24 Mar
  12. ForumIssue #356913 Aug
  13. ForumIssue #396183 Sep
  14. ForumIssue #401297 Sep
  15. ForumIssue #4065411 Sep
  16. ForumIssue #4072811 Sep
  17. ForumLitelm: LiteLLM Without the Bloat11 Sep
  18. ForumIssue #4149216 Sep
  19. OtherAI Gateway Latency Benchmarks9 Sep
  20. OtherLLM Gateways in Production: LiteLLM vs Portkey vs Bifrostread 18 Sep
  21. OtherAI Gateway Benchmark: Bifrost vs LiteLLMread 18 Sep

MethodOn 18 September 2026 we read the project's README, licence file, spend-tracking page, release list and its own benchmark post, two editorial assessments from 15 May 2025 and 15 April 2026, one news report from 24 March 2026, one repository API call and one issue-search API call per theme for the totals, five open issues and two developer threads as ordinary pages, and synthesised them with AI; no gateway was installed, no request was served and nothing was tested.

Dates without a year are 2026.

The review as text

The same review as one piece of writing · 7 min read

LiteLLM is BerriAI's open-source AI gateway: a Python SDK and a self-hosted proxy that put more than 100 model providers behind one OpenAI-shaped call, with virtual keys, budgets and spend logs on top. The release read here is v1.101.0, published on 15 September 2026; every figure below is as its source gave it on 18 September 2026.

We read the README, the licence file, the spend-tracking page, the release list and the maker's benchmark post; one GitHub search API call per complaint for the open-issue totals; five issues and two developer threads as ordinary pages; the Thoughtworks Radar entry of 15 April 2026; InfoWorld's assessment of 15 May 2025; and BleepingComputer's report of 24 March 2026.

Consensus

LiteLLM is described as the broadest of the self-hosted gateways, with its centre of gravity in governance rather than speed. Both editorial assessments describe the same core: one OpenAI-shaped call over many providers, with the value in virtual keys, budgets and spend tracking. Three comparisons narrow the difference to latency.

Confidence is strongest on the surface and weakest on speed: every performance figure we could reach comes from a test with a product on one side. This is a reading of v1.101.0 on 18 September 2026, when GitHub returned 59,017 stars, 11,538 forks and 5,201 open issues.

Recurring strengths

One shaped call over many providers. Both editorial assessments and three comparisons name this as the reason it gets adopted. InfoWorld describes a universal adapter for provider APIs; Thoughtworks says its teams use it as a standard place to put governance, tracing and key management.

The governance surface. The same assessments and comparisons put the value in keys scoped per team or app, budgets and rate limits on those keys, request logs and spend tracking, which the README lists as shipped.

It ships often. GitHub's release list shows v1.101.0 on 15 September 2026 and a push on 17 September, so the release line resumed after the March pause.

A licence split stated in the open. Outside the enterprise directory the licence file is MIT, and the licence inside it is named on the first line rather than in a footnote.

Recurring complaints

Provider behaviour does not fit the OpenAI shape. One GitHub search API call returns 102 open issues with the word streaming in the title on 18 September 2026 (GitHub, one call). Issue #40654, dated 11 September 2026, describes reasoning text dropped when a Responses-shaped answer is bridged to Chat Completions; #41492, dated 16 September 2026, describes Gemini grounding metadata dropped on streaming responses. Thoughtworks reports the same fault line from outside: teams relying on a provider's own capabilities pass provider-specific parameters anyway, and its entry says drop_params discards unsupported parameters silently. The 102 is a count of issues, not a rate.

Cost and spend accounting takes work. A second search API call returns 71 open issues with the word cost in the title that day. #39618, dated 3 September 2026, describes the cost path failing on missing token counts; #35691, dated 3 August 2026, describes spend logs recording zero for custom models outside the built-in cost map; #40728, dated 11 September 2026, describes an Azure model router path with no cost tracking. The spend-tracking page says it tracks spend for all known models and documents a hole of its own: an end user can send a user id in the body and have the spend land on that name instead of the key.

It is not light, whatever the name says. A third call returns 18 open issues with the word memory in the title, and the Hacker News thread of 11 September 2026 (176 points, 63 comments) is about a fork built by cutting features out, with the comments naming cost tracking and streaming among the cuts as the reason people run the gateway. The maker's benchmark post of 22 July 2026 reports a p99 of 0.7 ms of added latency and 21.8 MB of peak memory for its Rust gateway against 257.7 ms and 329.5 MB for the legacy Python proxy, on a local mock upstream.

Two compromised releases. BleepingComputer reported on 24 March 2026 that versions 1.82.7 and 1.82.8 went to PyPI carrying a credential stealer, citing research by Endor Labs and reporting roughly 500,000 exfiltration events it could not confirm independently. The maintainer's statement that day says the Docker proxy path pins its dependencies and was not affected. That characterisation is the project's own.

Where reviewers split

Performance, claimed against measured. One side is the maker's post, with sub-millisecond overhead on a mock upstream. The other is DeepInspect's piece of 9 September 2026, which argues a mock upstream deletes the variable that dominates real requests, and the Medium write-up, which measures time to first byte rising from 35.49 ms to 163.02 ms under load. Our read is that the two sides measure different things.

How far the incident reached. One side is the maintainer's account, that the pinned Docker path was untouched. The other is BleepingComputer's report of roughly 500,000 exfiltration events, which it says it could not confirm independently. So the exposure turns on how the gateway was installed, and none of the sources we read settled it.

Whether the weight is worth paying. One side is the September 2026 thread, where the weight and the feature list are the complaint. The other is the Medium benchmark's summary, which puts LiteLLM ahead on provider, endpoint and administration coverage while it loses on latency.

Who it suits

You already hold provider contracts of your own and want one OpenAI-shaped endpoint in front of them. Both editorial assessments describe that as the case it is built for.

You want per-key budgets, scoped keys and request logs you run yourself, and you will run a database and a cache to get them.

You will pin versions and read release notes. The two compromised releases turned on what an unpinned install could pick up.

You need a provider's own capabilities and can configure them explicitly. Thoughtworks reports they do not survive the translation alone.

This is a reading of published reviews, not medical, financial or legal advice. For a decision about your health, your money or your rights, a qualified professional is the right next step, and not a review.

Who should pass

You are choosing it for raw speed on the strength of a benchmark. No test without a product in the comparison turned up.

You expect a provider's own response fields to arrive unchanged through a streaming bridge. That is what 102 open issues with streaming in the title describe.

You cannot keep the install pinned to known-good versions. Two releases went to PyPI carrying a credential stealer on 24 March 2026.

You want the smallest possible dependency tree. Eighteen open issues name memory in the title, and the thread of 11 September 2026 is a fork built by cutting features out.

One verdict a week: the most useful Master Review we finished, the complaint that kept appearing, and who should skip it. Get the weekly verdict.

Sources

On 18 September 2026 we read the project's README, licence file, spend-tracking page, release list and its own benchmark post, two editorial assessments from 15 May 2025 and 15 April 2026, one news report from 24 March 2026, one repository API call and one issue-search API call per theme for the totals, five open issues and two developer threads as ordinary pages, and synthesised them with AI; no gateway was installed, no request was served and nothing was tested.

One verdict a week.

Every week, the most useful Master Review we finished: what the internet agrees on, the complaint that kept appearing, and who should skip it.

By subscribing you agree to receive one weekly email from Review Machine. You can unsubscribe at any time.