KPI

Accuracy: AI benchmark

Share of outputs that are correct on a checked sample.

How to measure it

Human review of a random sample of outputs against a written standard.

Across the library

Median 91% across 21 deployments, reported range 42% to 100%, plus 1 "up to" value not pooled.

Higher is better. Unit: percent.

Reported values by use case

AI for IT incident triage and root cause analysis (AIOps)

Use cases that should track accuracy