AI red-team evaluation safety certification has formal limits that a new arXiv paper makes precise: OpenAI's sandbox escape, ...