No Single Safeguard
This week: Reliability theory for AI Control, auditing sabotage bench, autobenchmark, optimal timing for superintelligence, serious AI product.
This week: Reliability theory for AI Control, auditing sabotage bench, autobenchmark, optimal timing for superintelligence, serious AI product.