Back to explore

Jev vs Gemini: Test di gestione dell'output strutturato

L'autore ha ampliato il batch a circa 300 domande per chiamata. Jev ha risposto correttamente, mentre Gemini ha rifiutato tutte le richieste prima di generare qualsiasi cosa. L'input era di soli ~55k token in una finestra di contesto di ~1M, indicando un problema nel modo in cui Gemini impone l'output strutturato.

5/7 Then I widened the batch to roughly 300 questions per call. ✅ Jev answered. ❌ Gemini rejected all 300 test requests before generating anything. The input was only ~55k tokens in a ~1M-token context window. The problem was how Gemini enforces structured output using a

Jev vs Gemini: Test di gestione dell'output strutturato 1
· 0 likesOpen on X