Back to explore

Jev security questionnaire response classification evaluation and tuning

The post discusses Jev's performance in categorizing security questionnaire responses (implemented, N/A, not implemented), comparing with Claude, noting that Jev needed tuning for more accurate results.

#Jev security questionnaire response categorization (implemented, N/A, not implemented) given the question+response then a ground truth eval. - Claude initially performed better with little direction - Jev questions needed tuning to provide a more accurate response. Once tuned,

Jev security questionnaire response classification evaluation and tuning 1
· 0 likesOpen on X