Agent Evaluation Benchmark
A confidence-aware evaluation framework for openclaw-class personal assistants, built to measure more than demo performance.
Applied statistical confidence, judge agreement, coverage, and cross-run stability; a separate expansion added 40 RPA, IM, cross-system, and email tasks.