V&V takes on “Pacing the frontier”

[Cross-posted to LW. These short takes try to put a verification-and-validation slant on AI-safety / alignment topics – they are not full treatments. I co-originated coverage-driven verification (CDV), and spent several decades doing verification of chips and AVs. See intro post for background.] Summary: The latest frontier model incidents resulted in the “Pacing the frontier” letter and subsequent … More V&V takes on “Pacing the frontier”

V&V takes on OpenAI’s long-horizon incidents

[Cross-posted to LW. These short takes try to put a verification-and-validation slant on AI-safety / alignment topics – they are not full treatments. I co-originated coverage-driven verification (CDV), and spent several decades doing verification of chips and AVs. See intro post for background.] On July 20 and 21, OpenAI published two unusually candid incident reports: … More V&V takes on OpenAI’s long-horizon incidents

When is misalignment just a bug? 

[LW linkpost is here] Introduction and epistemic status: This is the first post in a planned series, “Alignment as a verification problem”. I co-originated coverage-driven verification (CDV), which became the standard methodology for chip verification and is heavily used in AV safety. Back in 2015 I wrote that verifying “Friendly AI” would be our biggest … More When is misalignment just a bug? 

Coverage-driven alignment – What ‘Teaching Claude Why’ can borrow from AV verification

Summary: This post suggests that alignment training could benefit from coverage-driven verification. Anthropic recently reported that teaching Claude alignment rules (via pretraining-style next-token learning on alignment-related stories) is more effective than relying primarily on RL-style behavioral shaping. Some AV developers reached a related conclusion, but in addition tend to use a systematic, coverage-driven methodology for … More Coverage-driven alignment – What ‘Teaching Claude Why’ can borrow from AV verification

Misc stuff: The verification gap, ML training and more

This post covers recent updates in machine learning, autonomous systems and verification. It has four sections: Automation / ML keep accelerating, but verification of automation / ML seems to lag behind HVC is coming, and I plan to attend (and even present) The idea of training an ML-based system using synthetic inputs (which I like) … More Misc stuff: The verification gap, ML training and more

Misc stuff: Robotics, system simulations, AI

Here is another one of those multi-topic summaries: Dipping my toes into autonomous robot verification As I discussed here, I am exploring working with Kerstin Eder et al. (of Bristol U) regarding autonomous robot verification. So I started looking into this. I will not take you into the gory details, gentle reader (especially since I … More Misc stuff: Robotics, system simulations, AI