The runtime reliability layer for AI agents.
-
Updated
Jul 24, 2026 - Python
The runtime reliability layer for AI agents.
Rage4j is a java library thats helps evaluate LLM's based on scientifically grounded metrics
Semantics aware quality evaluation of building 3D models: a learning approach
The PLATO tutoring language — meaning-matched response evaluation using single embedding space for semantic answer comparison in Rust
A simple AI chatbot testing framework
Reproducible evaluation of Arabic–English neural machine translation using lexical, character-based, and semantic quality metrics.
Work-in-progress research repository for a revised study comparing Transformer training strategies in Arabic–English machine translation.
SemEval-2025
Research prototype for meaning-preserving text transformation with semantic and quality evaluation.
Semantic Evaluation dataset for Uzbek language
SemEval2025
Add a description, image, and links to the semantic-evaluation topic page so that developers can more easily learn about it.
To associate your repository with the semantic-evaluation topic, visit your repo's landing page and select "manage topics."