6 views
-/https://github.com/berriai/litellm/issues/28422
GitHub · issue

#28422 [Bug]: Why is my locally deployed Qwen 14B always returning JSON-formatted responses, while the cloud API outputs normal conversational text? What happened?

  • State: open
  • Author: @Neohypery
  • Labels: bug

### Check for existing issues

- [x] I have searched the existing issues and checked that my issue is not a duplicate.

### What happened?

I deployed both a cloud model and a local Qwen14B model using LiteLLM. I noticed that when calling the local 14B model, the responses are always in JSON format. I tried disabling thinking mode, enabling streaming output, and modifying the config, but none of these changes worked.

Can anyone help explain how to properly solve this issue?

### Steps to Reproduce

* Changed reasoning: true → Probably unrelated, because other models are also running with false and work fine. * Added a system_prompt telling the model not to output JSON → Ineffective guess. * Enabled stream: true → Ineffective guess. * Modified SOUL.md → The problem was not related to this. * Created qwen3-nothink to remove thinking tags → Direction was partially correct, but repeated attempts caused errors and created duplicate models, wasting disk space. * Rebuilt the Modelfile multiple times → Re-downloaded the 9GB model every time. * Modified feishu-channel-rules → Problem was unrelated. Repeatedly deleted and recreated models based on guesses …

GitHub resolver

Import GitHub neighbors on demand. Results are saved as system ingests.

Refresh page
vote history (3 events)
#0 of 0 · 31d19h7m53s ago — entered · #import:https:::github.com:berriai:litellm post #1941
The right-hand change is a broad cross-layer feature involving configuration, date-boundary calculations, reporting queries, exports, UI consistency, timezone edge cases, and backward compatibility. The left-hand work is more likely a localized provider-behavior investigation and configuration or documentation fix.
30728 requires coordinated changes to guardrail failure semantics, multiple request/response pathways, streaming handling, and comprehensive regression coverage. 28422 appears primarily to require provider-specific diagnosis and configuration clarification, with uncertain or limited LiteLLM code scope.
#0 of 0 · 31d18h45m35s ago — current · #import:https:::github.com:berriai:litellm post #2315
The left task requires coordinated asynchronous streaming changes, per-deployment configuration, timing behavior, and regression testing across proxy integrations. The right task is primarily provider/model configuration diagnosis with comparatively limited product-code scope.
discussed in #import:https:::github.com:berriai:litellm

ranked child groups

no voted pairs yet in this scope

cli
src
spread
search