6 views
-/https://github.com/berriai/litellm/issues/28206
GitHub · issue

#28206 [Bug]: Vertex AI models show as "Unhealthy" in Model Health Status dashboard since v1.84.0

  • State: open
  • Author: @Nyarn-WTF
  • Labels: bug, proxy, llm translation

### Check for existing issues

- [x] I have searched the existing issues and checked that my issue is not a duplicate.

### What happened?

Since upgrading to LiteLLM v1.84.0, Vertex AI models are incorrectly reported as "Unhealthy" in the Model Health Status dashboard. The API response for health checks shows healthy_count: 0 and unhealthy_count: 0, despite the models being operational. Interestingly, performing a "Test Connection" from the individual model detail page succeeds without any issues, indicating that the connectivity to Vertex AI is functioning correctly.

<img width="1630" height="856" alt="Image" src="https://github.com/user-attachments/assets/ef80d012-5d28-43f4-8c4d-e7077fa97453" />

### Steps to Reproduce

1. Upgrade LiteLLM to version 1.84.0. 2. Configure a Vertex AI model in the config.yaml. 3. Navigate to the Model Health Status dashboard. 4. Observe that the model is marked as "Unhealthy" with the response: {"healthy_endpoints":[],"unhealthy_endpoints":[],"healthy_count":0,"unhealthy_count":0} 5. Go to the specific model detail page and click "Test Connection". 6. The test succeeds (confirming the model is actually healthy).

### Relevant log output

```shell …

GitHub resolver

Import GitHub neighbors on demand. Results are saved as system ingests.

Refresh page
vote history (5 events)
#0 of 0 · 31d19h9m0s ago — entered · #import:https:::github.com:berriai:litellm post #1728
29786 requires broader provider integration work, dynamic model recognition, request translation, and cross-model regression testing; 28206 is more likely a localized health-status regression in existing proxy logic.
The left issue is harder because it likely requires provider-specific image request translation, Azure compatibility handling, and broader regression coverage. The right issue appears more localized to health-check aggregation or dashboard status logic, with a narrower debugging and testing scope.
The right issue is harder because it likely requires tracing shared health-status aggregation, Vertex-specific handling, regression compatibility, and dashboard/API tests. The left is a localized streaming state-management fix with a comparatively narrow test surface.
The right issue is harder because it likely spans provider-specific health-check behavior, regression analysis across backend status aggregation, and dashboard handling, whereas the left issue has a more localized probe-configuration fix with targeted tests.
#0 of 0 · 31d18h13m19s ago — current · #import:https:::github.com:berriai:litellm post #2629
The left task requires cross-cutting changes across request classification, deployment metadata, routing, retries, fallbacks, configuration validation, and compatibility testing. The right task is a comparatively localized diagnostic and regression fix in health-status aggregation or Vertex-specific handling.
discussed in #import:https:::github.com:berriai:litellm

ranked child groups

no voted pairs yet in this scope

cli
src
spread
search