OpenAI Expands Outside Safety Reviews Into Model Training
OpenAI plans to expand independent safety assessments into model training and evaluation, giving outside groups earlier access to test high-risk AI systems. The post OpenAI Expands Outside Safety Reviews Into Model Training appeared first on TechRepublic .
OpenAI is broadening its independent safety reviews to encompass model training and evaluation, enabling external groups to scrutinize high-risk AI systems earlier in the development process. This shift moves beyond the company's previous reliance on safety assessments solely before model launch dates. OpenAI's Lama Ahmad, who oversees partnerships with outside safety experts, explained that the company has historically engaged third parties shortly before deployment.
However, with escalating stakes, OpenAI now intends to involve independent evaluators during training, evaluation, and deployment phases.
Ahmad outlined four key areas where independent scrutiny will focus: validating the evidence supporting OpenAI's safety claims, testing critical safeguards against adversarial attacks like jailbreaks, evaluating defenses against high-risk capabilities such as cybersecurity threats and biological misuse, and verifying that alignment evaluations can identify significant misalignments. Some assessments could span weeks to months, allowing ongoing analysis of safety claims rather than serving as a mere pre-launch verification.
To ensure transparent yet secure assessments, OpenAI stipulates that evaluators must align on predefined claims, receive appropriately proportionate access to data and systems, and disclose any conflicts of interest. Sensitive information may necessitate assessments occurring within company-managed environments, albeit this tight restriction might curb the independence of outside researchers testing the company's claims.
While conducting assessments during training offers the advantage of identifying vulnerabilities before deployment, it does not guarantee robust oversight. The efficacy of this independent scrutiny hinges on the evaluators' qualifications, the depth of system access granted, and the scope of claims cleared for testing. OpenAI has not yet identified official assessment partners or defined public access guidelines, leaving operational details uncertain.
Nevertheless, this initiative holds the potential to uncover alignment, security, and misuse risks early in the development lifecycle, enabling OpenAI to preemptively address issues before systems are deployed. However, the extent to which this model will genuinely enhance safety protections remains to be seen, as full transparency may be constrained to safeguard proprietary systems and sensitive vulnerabilities.
Written by urgent.news from TechRepublic's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.