stuntdouble: Shadow proxy evaluation for Jev
A zero-dependency Node proxy that shadows TypeSafe Jev with local models (Kev, Laya) on live traffic and produces a report on whether you can swap.
OPEN SOURCE · CURATED REPOSITORIES
Curated Jev SDKs, tools, demos, integrations, benchmarks, and research projects. Repository metrics come from GitHub; titles, summaries, and categories are editorial.
Ranked by GitHub Stars
Ranked by recent Star growth and update activity
A zero-dependency Node proxy that shadows TypeSafe Jev with local models (Kev, Laya) on live traffic and produces a report on whether you can swap.
A CityFlow-based traffic signal control framework that evaluates TypeSafe Jev and OpenAI-compatible LLM controllers alongside rule-based and reinforcement-learning baselines through the same simulation pipeline.
A CLI that reads local coding agent session logs and scores them using Jev, TypeSafe AI's System One model, producing interactive HTML reports and terminal dashboards.
Calls TypeSafe AI's Jev/System One model via the official SDK, compares it with an LLM stand-in across 24 operational decisions, and commits reproducible run results and evaluation reports.
A CLI/evaluation tool for TypeSafe Jev (System One): measures thresholds, calibrates confidence, and checks drift on your own data, with optional LLM-teacher labels. Contains a no-key demo simulator and live integration with the Jev API.
A pre-registered audit of TypeSafe AI's Jev on Spanish — accuracy, calibration, and token cost — plus a CLI to run the same comparison on your own labelled data.
An open-source local alternative to TypeSafe's Jev System One model, providing typed decisions with calibrated confidence. Includes a cross-system benchmark against Jev and other local models.
Three measured experiments on RAG hallucination: quote-checking, TypeSafe's Jev, and IBM's STAIR. 850+ graded questions, raw responses included.
Journey Evals is an evaluation tool that drives a real browser or LangGraph agent through declared user journeys and verifies outcomes with backend evidence; it uses TypeSafe's Jev for operation and element selection.
Open-source content analytics and LLM-as-judge harness that rates posts with calibrated judgments from TypeSafe Jev (System One) and tests them against real engagement data.
Code and raw data behind Lunar Terminal: Robocode Tank Royale bots, GapFlap, a token-metering proxy, and a 327-decision randomised trial of TypeSafe Jev.
An agent skill for Claude Code, Codex and Cursor that scans a repo or idea, checks Jev fit and failure modes, estimates cost/latency, and runs a live pilot against the TypeSafe API.