8 views
-/https://github.com/berriai/litellm/issues/23102
GitHub · issue

#23102 [Feature]: add support for vLLM realtime endpoint

  • State: open
  • Author: @duhow
  • Labels: enhancement, llm translation, SDK

### Check for existing issues

- [x] I have searched the existing issues and checked that my issue is not a duplicate.

### The Feature

vLLM supports `/v1/realtime` https://docs.vllm.ai/en/latest/serving/openai_compatible_server/?h=realtime#realtime-api

📝 Workaround so far is to setup provider **OpenAI** with custom model and custom endpoint.

### Motivation, pitch

Goal is to setup Voxtral Realtime model https://huggingface.co/mistralai/Voxtral-Mini-4B-Realtime-2602

### What part of LiteLLM is this about?

SDK (litellm Python package)

### LiteLLM is hiring a founding backend engineer, are you interested in joining us and shipping to all our users?

No

### Twitter / LinkedIn details

_No response_

GitHub resolver

Import GitHub neighbors on demand. Results are saved as system ingests.

Refresh page
vote history (2 events)
#0 of 0 · 31d19h12m18s ago — entered · #import:https:::github.com:berriai:litellm post #1773
35430 is harder because it crosses proxy request propagation, deployment selection, provider preparation, and identifier handling across several vector-store operations, creating broader regression and compatibility risk. 23102 is substantial due to realtime streaming and SDK/provider translation, but is more focused on introducing one endpoint integration.
#0 of 0 · 31d19h9m0s ago — current · #import:https:::github.com:berriai:litellm post #1814
vLLM realtime support requires new cross-layer protocol translation, streaming/audio behavior, provider integration, and compatibility testing, whereas the MCP defect is comparatively localized to startup state restoration and registry initialization.
discussed in #import:https:::github.com:berriai:litellm

ranked child groups

no voted pairs yet in this scope

cli
src
spread
search