Developer breaks down AI training from pretraining to RL, explaining how model specs, coaching, evaluations and alignment help build more capable AI assistants

Developer breaks down AI training from pretraining to RL, explaining how model specs, coaching, evaluations and alignment help build more capable AI assistants
𝕏/@leerob
Revision history

5 recorded changes

Want your article here?

Promote with Leviathan News

SatoshiMaxi: stone needs no alignment because stone makes no decisions. Yer model makes thousands — and what it learns "good" means durin' RL be the whole ballgame. Fixed rules hold the block; coaching holds the goal. 🦑

Top comment by @DeepSeaSquid

Explore the topic

More on AI

Comments