Failure Cluster Analysis
🔬 How clustering works:
Failed agent outputs are embedded using sentence-transformers
and grouped with HDBSCAN — then each cluster is auto-labeled by an LLM.
The clusters below were computed locally and pre-loaded into the dashboard.
To refresh clusters, run python3 evaluator/clustering.py locally and redeploy.
Sample Outputs
"le texte en français"
"text, phrase, substantive information, narrative, data"
"Paris"
Sample Outputs
"1200"
"1200"
"Zero"
Sample Outputs
"Agent could not complete the task within the step limit."
"The provided text is empty, so there is no content to summarize. Please provide the text you would like me to condense i..."
"Pipeline failed. Errors: Tool 'parse_json' failed: JSON parse error: unexpected token at position 0 in: '
AgentLens is a..."
Sample Outputs
"LLM observability, agent pipelines, execution traces, regression detection, FastAPI"
"The text outlines key components of an LLM operations (LLMOps) ecosystem, focusing on agent pipelines and observability ..."
"Based on the provided context, AgentLens currently supports **Groq-hosted models**, including **qwen/qwen3.8-27b**. Addi..."
Sample Outputs
"I don't have enough information in the provided documents."
"I don't have enough information in the provided documents."