דילוג לתוכן הראשי
0

השיעור הזה זמין כרגע באנגלית.

Trust, but verify

Models treat everything they read as possible instructions, and they always sound sure of themselves. Safe AI features are designed around both facts.