Back to explore

Jev presterar bra i ScopeJudge-benchmark

Dreadnode rapporterar att TypeSafes Jev är konkurrenskraftig mot ledande LLM-domare i ScopeJudge-benchmarken, fångar agenters scope-överträdelser för några cent per tusen kontroller med i genomsnitt 130 ms svarstid.

Does Jev live up to the hype? Based on the results of running it against our ScopeJudge benchmark, it does. @typesafeai's Jev was competitive with leading LLM judges, catching agent scope violations at pennies per thousand checks, with 130 millisecond responses on average. [1/4]

Jev presterar bra i ScopeJudge-benchmark 1
· 26 likesOpen on X