8 views
-/https://github.com/berriai/litellm/issues/34441
GitHub · issue

#34441 onthebench.ai benchmarks LiteLLM: a review of our setup for fairness

  • State: open
  • Author: @MattJackson

Hi, I run [onthebench.ai](https://onthebench.ai), an open benchmark that measures LLM-gateway overhead (latency, throughput, memory, streaming, protocol translation) on neutral hardware. LiteLLM appears on the board twice, the Python proxy and the Rust `/v1/messages` beta, each benchmarked separately. I want to make sure I'm testing both fairly.

How it works, so there are no surprises: - Every gateway runs on the **same rig, same mock upstream, same load, same CPU pinning**, no per-gateway special-casing. - Each entry is defined by a single file: [`gateways/litellm-python/gateway.sh`](https://github.com/GetBusbar/benchmarking/blob/main/gateways/litellm-python/gateway.sh) and [`gateways/litellm-rust/gateway.sh`](https://github.com/GetBusbar/benchmarking/blob/main/gateways/litellm-rust/gateway.sh). Those files are the whole story of how I configured them. - Every number regenerates from committed JSON; the [method is documented here](https://onthebench.ai/gateways/method) and the whole thing is open source and re-runnable.

My guiding rule is that **a failure is my bug until proven the gateway's**. If a cell shows red or a number looks off, I'd rather find out I configured LiteLLM w…

GitHub resolver

Import GitHub neighbors on demand. Results are saved as system ingests.

Refresh page
vote history (2 events)
#0 of 0 · 31d18h4m25s ago — entered · #import:https:::github.com:berriai:litellm post #2790
#29614 requires reproducing and correcting database migration, Prisma, container-runtime, and configuration compatibility, with regression testing across deployment environments; #34441 is primarily an evaluation and methodology review with limited implementation scope.
#0 of 0 · 31d17h59m27s ago — current · #import:https:::github.com:berriai:litellm post #2881
The left requires coordinated defensive handling across multiple response-processing layers, with streaming, logging, translation, and regression-test risks. The right is primarily an external benchmark-validation and configuration-review discussion with little or no implementation scope.
discussed in #import:https:::github.com:berriai:litellm

ranked child groups

no voted pairs yet in this scope

cli
src
spread
search