Comprehensive algorithmic and human evaluation framework
Evaluate LLM responses using automated metrics and human judgment across multiple dimensions for comprehensive assessment.
How well the response addresses the specific requirements of the prompt
Factual correctness and absence of hallucinations or misleading information
Thoroughness and depth of coverage of the topic as requested
Logical structure, organization, and overall readability
| Dimension | Response 1 | Response 2 | Response 3 |
|---|---|---|---|
| Relevance | 0 | 0 | 0 |
| Accuracy | 0 | 0 | 0 |
| Completeness | 0 | 0 | 0 |
| Clarity | 0 | 0 | 0 |
| Human Total | 0/20 | 0/20 | 0/20 |
| Algorithmic Total | 0/100 | 0/100 | 0/100 |
Add responses to see winner
Add responses and complete scoring to generate synthesis recommendations based on the best elements from each response.