lecture July 31, 2026 18:04 YouTube

Why Jev — Diogo Almeida, TypeSafe AI

@aiDotEngineer

AIRLHFAlignmentAutomation

Almeida gives the argument behind Jev weeks before the model ships, while TypeSafe is, in his words, “still kind of stealthy.” The case is that RLHF optimises for human preference, that a reward model finds visible uncertainty far easier to punish than incorrectness is to detect, and that hallucination therefore follows from the objective rather than from any missing patch — his illustration is ChatGPT calling a recording of fart noises “a very eerie vibe atmosphere piece.” He splits post-training into three branches by what each targets, RLHF for preference, RLVR for correctness, and TypeSafe’s RLCD for calibrated decisions, and defends the part he is not attacking: pre-training is “phenomenal,” and the problem is “how we unearth it.” Worth watching for what he declines to answer, too — asked directly why not simply train a classifier head on top of pre-training, he says it’s complicated and moves on.

Theme
Language
Support
© funclosure 2025