用 Jev 分类模型逐步选择合法 token 编写 C++
作者用 TypeSafe 的 Jev 决策/分类模型做实验:每步只提供语法合法的候选 token,让模型选择,从而生成 C++ 代码。
Dorian Smiley 发布了状态机编程基准:Gemini Flash Lite 3.1 总体 75.6%,Jev 总体 72.2%,但在典型任务上达 100%,速度快约 6 倍,失误主要集中在新的组合场景。
New state machine programming benchmark: Gemini Flash Lite 3.1: 75.6% overall, 66.7% generalization. Jev: 72.2% overall, 61.4% generalization, but 100% on canonical tasks and ~6× faster. Jev’s misses are concentrated in specific novel compositions, not broad task failure. It is