OPEN SOURCE · CURATED REPOSITORIES

Jev GitHub Projects.

Curated Jev SDKs, tools, demos, integrations, benchmarks, and research projects. Repository metrics come from GitHub; titles, summaries, and categories are editorial.

750 projects
13 projects
sysadarshsysadarsh/zerosweep
BenchmarkDeveloper toolsTypeScript

ZeroSweep: TypeSafe Jev Triage Engine & Benchmark

An autonomous triage engine and benchmark powered by TypeSafe AI's Jev System One model, demonstrating high-speed email classification with calibrated confidence.

200
JevalsJevals/jevals-data
BenchmarkScientific research

Jevals Data: Independent benchmark for TypeSafe Jev

Independent benchmark data comparing TypeSafe AI's Jev (System One model) against LLMs on decision score, accuracy, calibration, cost, and latency for typed tasks.

100CC-BY-4.0
arifulislamatarifulislamat/jev-benchmark
BenchmarkScientific researchTypeScript

Jev Benchmark

Benchmarks Jev (typesafe/jev-1.13) via OpenRouter's decisions endpoint, measuring typed decision cost, latency, and accuracy on 100 support tickets against Claude, GPT, and Gemini.

000MIT
Sunwood-ai-labsSunwood-ai-labs/jevdash
BenchmarkDeveloper toolsPython

JevDash: System One

JevDash is a Pygame-based 2D platformer benchmark that evaluates TypeSafe AI's Jev model in real-time control scenarios, with Vercel AI Gateway integration.

000MIT
planstack-aiplanstack-ai/jev-tetris-benchmark
BenchmarkDeveloper toolsTypeScript

Jev Tetris Benchmark

A reproducible Tetris decision benchmark that calls TypeSafe Jev / System One and Claude Haiku on identical boards, demonstrating Jev Choice integration via Vercel AI Gateway or the TypeSafe API.

000MIT
robokrunchrobokrunch/jev-physical-ai
BenchmarkScientific researchPython

Jev Physical AI Benchmark

Real-world benchmark of TypeSafe's Jev for robot fleets and edge hardware, with 300 measured API calls, cost analysis, and reproducible scripts.

200MIT
clduab11clduab11/jev-test
BenchmarkScientific researchPython

Jev-Test: Pre-registered benchmark for TypeSafe Jev as decision model

Pre-registered benchmark testing whether TypeSafe Jev's confidence scores reliably decide search, evidence, and citation support for a 2B local model; includes harness, audit tool, and open end-to-end results.

100MIT
emretheusemretheus/jev-rag-benchmark
BenchmarkScientific researchPython

Jev RAG Benchmark

A free RAG reranking and decision-layer benchmark for TypeSafe Jev 1.13, using frozen candidate pools, paired bootstrap CIs, and calibration, comparing against NVIDIA cross-encoder and no-reranking baselines.

000MIT
PsiACEPsiACE/dohnuts
BenchmarkScientific researchPython

Dohnuts: Small multimodal decision model compared with Jev

This project builds small multimodal models for direct decisions and benchmarks them against TypeSafe Jev and Laya in its model card.

2430Apache-2.0
instax-duttainstax-dutta/sysone-bench
BenchmarkScientific researchPython

sysone-bench: Jev vs Laya benchmark

First independent head-to-head benchmark of System One decision models Laya and Jev on byte-identical inputs, reporting accuracy, ECE, gating, and latency.

410MIT
EnesDemir143EnesDemir143/jev-laya-benchmark
BenchmarkDeveloper toolsPython

Jev vs Laya Benchmark

Local benchmark comparing TypeSafe Jev and Laya-MLX for structured issue classification. Uses the Jev API with Noul scores.

100
DililianxiceDililianxice/jev-inner-speech-bci
BenchmarkScientific researchPython

Jev × Inner-Speech BCI Benchmark

Reproducible benchmark evaluating Jev as a semantic prior for intracortical inner-speech BCI decoding. Reports Jev's input-efficiency signal in synthetic interactions and neural decoding limits on public data.

100MIT