Developers of general-purpose AI should run safety evaluations that scale with model capability, for example assessing cyber-offensive capabilities and vulnerabilities, testing for chemical, biological, radiological and nuclear information risks, evaluating behaviour beyond intended uses, testing for jailbreaking and prompt manipulation, data privacy risks, or comprehensive red teaming.
The graph holds this control, the 0 it maps to, and the evidence behind each claim, over MCP and REST.