AI Safety
9 items tagged with "ai-safety"
Best Practices2
OpenAI Safety & Alignment Best Practices
Mitigation strategies (RLHF, red-teaming, tiered access) for large language model deployment.
AI Red Teaming
AI red teaming is structured adversarial testing of AI systems to find harmful, biased, or insecure behavior before attackers or real users do, using crafted attacks and probes.
Models3
GPT-6 Astra
OpenAI's GPT-6 Astra is described as its most intelligent and aligned broadly deployed model, with state-of-the-art capabilities across computer use, coding, cybersecurity, and science. It is the first OpenAI model in the provided data to reach the Critical cybersecurity capability threshold under the Preparedness Framework.
Astra
Astra is an OpenAI frontier model announced as the first OpenAI model to meet the Critical cybersecurity capability threshold under the Preparedness Framework. The announcement emphasizes stronger safeguards for release.
GPT-5.6 Sol
GPT-5.6 Sol is a next-generation OpenAI model previewed with stronger capabilities in coding, science, and cybersecurity, paired with OpenAI's most advanced safety stack.
Regulations3
Executive Order 14110 on Safe, Secure, and Trustworthy Development and Use of Artificial Intelligence
US federal executive order directing agencies to set safety, security, and rights standards for AI development and deployment across government and industry.
Artificial Intelligence and Data Act
Proposed Canadian federal law (part of Bill C-27) regulating high-impact AI systems for safety and non-discrimination.
EU Machinery Regulation (Regulation (EU) 2023/1230)
EU regulation setting health, safety, and cybersecurity requirements for machinery and related products, including digital and AI-enabled systems.