organizations, and more scalable sharing of cyber threat intelligence that can prevent the catastrophic scaling of attacks. However, new incentives and investments are necessary to create a coordination layer that provides herd immunity against AI-enabled attacks.
There may be opportunities to mitigate near-term risks and create time for defensive adaptation. Approaches such as controlled access to U.S. leading-edge AI models, export controls on critical hardware, or limits on model distillation—whether through voluntary industry self-governance or policy measures—may slow the diffusion of the most powerful offensive tools. In parallel, regulatory tools can play an important role in shaping incentives. Revisiting mechanisms such as financial disclosure requirements may help make cybersecurity risks more visible, while new regulatory approaches may be warranted in critical sectors where the consequences of failure are particularly high. However, these measures are unlikely to be comprehensive or durable in isolation and should be understood as mechanisms for buying time in a rapidly evolving global ecosystem.
Longer-term defensive advantages will require sustained investment and coordination. With investment across technical, architectural, and institutional dimensions, frontier AI could enable a fundamentally stronger approach to cybersecurity that may shift the advantage to the defenders. High-fidelity digital twins of critical systems can support continuous assessment and testing, particularly in high-consequence environments. At the level of architecture, separating data collection from reasoning, as well as enabling shared knowledge bases, will allow defensive systems to improve alongside advances in AI capabilities. AI may also help strengthen the underlying cybersecurity foundation by enabling more effective secure-by-design development practices and reducing reliance on purely reactive approaches to vulnerability remediation. At the institutional level, governance arrangements that incorporate academic, governmental, and nonprofit participation will be necessary to ensure independent evaluation, transparency, and sustained investment in foundational research. Public–private partnerships will be central to aligning incentives, sharing risk-relevant information, and developing scalable approaches to cybersecurity.
The short-term outlook is therefore concerning, while the longer-term outlook is cautiously optimistic. The central task for policy makers and practitioners is to shorten the interval between these two regimes—mitigating near-term risks while accelerating the deployment of more adaptive, scalable, and resilient defensive capabilities. The case for acting now is straightforward: declining to adopt AI-enabled cybersecurity defense would leave mission-critical infrastructure exposed to autonomous, intelligent adversaries operating at a scale beyond human response. Acting responsibly, but acting, is the only viable path.
AI Security Institute. n.d. "Rigorous AI Research to Enable Advanced AI Governance." Department for Science, Innovation, and Technology. https://www.aisi.gov.uk (accessed May 4, 2026).
Anthropic. 2026. Project Glasswing: Securing Critical Software for the AI Era. https://www.anthropic.com/glasswing (accessed May 4, 2026).
Castro, Sebastián R., Roberto Campbell, Nancy Lau, Octavio Villalobos, Jiaqi Duan, and Alvaro A. Cardenas. 2025. "Large Language Models are Autonomous Cyber Defenders." In 2025 IEEE Conference on Artificial Intelligence (CAI), 1125–32. Santa Clara, CA: IEEE. https://doi.org/10.1109/CAI64502.2025.00195.
Conger, Kate. 2026. "One Job That Is Growing in the A.I. Era? Cybersecurity Experts." The New York Times, May 24, 2026, updated May 25, 2026. https://www.nytimes.com/2026/05/24/technology/ai-cybersecurity-jobs.html.
DARPA (Defense Advanced Research Projects Agency). 2025. Artificial Intelligence Cyber Challenge. https://www.aicyberchallenge.com (accessed April 17, 2026).
de Gregorio, Alfonso. 2025. "Mitigating Cyber Risk in the Age of Open-Weight LLMs: Policy Gaps and Technical Realities." arXiv preprint arXiv:2505.17109.
Eastland, Maggie, and Shirin Ghaffary. 2026. "Google, Microsoft to Give US Agency Early Access to AI Models." Bloomberg, May 5, 2026. https://www.bloomberg.com/news/articles/2026-05-05/ai-firms-agree-to-give-us-early-access-to-evaluate-their-models.
Folkerts, Linus, Will Payne, Simon Inman, Philippos Giavridis, Joe Skinner, Sam Deverett, James Aung, et al. 2026. "Measuring AI Agents' Progress on Multi-Step Cyber Attack Scenarios." arXiv preprint arXiv:2603.11214.
Future of Life Institute. 2026. The EU Artificial Intelligence Act. https://artificialintelligenceact.eu (accessed May 4, 2026).
Google. 2025. OSS-Fuzz README. GitHub Repository. https://github.com/google/oss-fuzz (accessed April 17, 2026).
Hicks, Mike, and Steve Lipner. 2026. "Beyond Penetrate-and-Patch: Why AI's Greatest Security Contribution Will Be Secure-by-Design." mhicks.mehttps://mhicks.me/blog/beyond-penetrate-and-patch/, May 5, 2026. (accessed May 26, 2026).
Holley, Bobby. "The Zero-Days Are Numbered." The Mozilla Blog. April 21, 2026. https://blog.mozilla.org/en/privacy-security/ai-security-zero-day-vulnerabilities/ (accessed May 27, 2026).
Holz, Thorsten. 2025. "Technical Perspective: Unsafe Code Still a Hurdle Copilot Must Clear." Communications of the ACM, January 21, 2025. https://cacm.acm.org/research-highlights/technical-perspective-unsafe-code-still-a-hurdle-copilot-must-clear/.
Howard, Michael. 2019. "Expert Tips for Finding Security Defects in Your Code." MSDN Magazine. Last modified October 21, 2019. https://learn.microsoft.com/en-us/archive/msdn-magazine/2003/november/expert-tips-for-finding-security-defects-in-your-code.
Jiang, Fengqing, Zhangchen Xu, Luyao Niu, Zhen Xiang, Bhaskar Ramasubramanian, Bo Li, and Radha Poovendran. 2024. "ArtPrompt: ASCII Art-Based Jailbreak Attacks Against Aligned LLMs." In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (1):15157–173.
Krašovec, Andraz, Gary Steri, Georgios Karopoulos, and Mirko Trapani. 2025. "Large Language Models for Cyber Threat Intelligence: Extracting MITRE with LLMs." In Availability, Reliability and Security. Lecture Notes in Computer Science 15995:80–9. Springer.
Li, Tao and Quanyan Zhu. 2025. "Agentic AI for Cyber Resilience: A New Security Paradigm and Its System-Theoretic Foundations." arXiv preprint arXiv:2512.22883.
Lockheed Martin. 2025. Cyber Kill Chain. https://www.lockheedmartin.com/en-us/capabilities/cyber/cyber-kill-chain.html (accessed May 4, 2026).
MITRE. 2026. ATT&CK Matrix for Enterprise. https://attack.mitre.org/ (accessed May 4, 2026).
Mohsin, Muhammad, Muhammad Umer, Ahsan Bilal, Zeeshan Memon, Muhammad Ibtsaam Qadir, Sagnik Bhattacharya, Hassan Rizwan, et al. 2025. "On the Fundamental Limits of LLMs at Scale." arXiv preprint arXiv:2511.12869.
Murgia, Madhumita and Jamie John. 2026. "Anthropic to expand Mythos access to more than 15 countries." Financial Times. June 2, 2026. https://www.ft.com/content/19ea9ed4-3dd0-43aa-9877-729d44e620e8?syn-25a6b1a6=1.
Naseem, Usman. 2026. "Mechanistic Interpretability for Large Language Model Alignment: Progress, Challenges, and Future Directions." arXiv preprint arXiv:2602.11180.
NASEM (National Academies of Sciences, Engineering, and Medicine). 2019. Implications of Artificial Intelligence for Cybersecurity: Proceedings of a Workshop. Washington, DC: The National Academies Press. https://doi.org/10.17226/25488.
NASEM. 2024. Foundational Research Gaps and Future Directions for Digital Twins. Washington, DC: The National Academies Press. https://doi.org/10.17226/26894
NASEM. 2025. Defense Software for a Contested Future: Agility, Assurance, and Incentives. Washington, DC: The National Academies Press. https://doi.org/10.17226/29129.
NIST (National Institute of Standards and Technology). 2023. AI Risk Management Framework. https://www.nist.gov/itl/ai-risk-management-framework (accessed May 4, 2026).
NIST. 2024. Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile (NIST AI 600-1). NIST Trustworthy and Responsible AI. Gaithersburg, MD: U.S. Department of Commerce. https://doi.org/10.6028/NIST.AI.600-1.
NIST. 2026. Cybersecurity Framework. https://www.nist.gov/cyberframework (accessed May 4, 2026).
NIST. n.d. Center for AI Standards and Innovation (CAISI). https://www.nist.gov/caisi (accessed May 4, 2026).
OpenAI. 2026. Introducing GPT‑5.5: A New Class of Intelligence for Real Work. https://openai.com/index/introducing-gpt-5-5 (accessed May 4, 2026).
Otto, Greg. 2026. "Researchers Say AI Just Broke Every Benchmark for Autonomous Cyber Capability." CyberScoop, May 13, 2026. https://cyberscoop.com/ai-autonomous-cyber-capability-benchmarks-broken-gpt5-claude-mythos/.
OWASP Gen AI Security Project. n.d. Agentic Security Initiative. https://genai.owasp.org/initiatives/agentic-security-initiative/ (accessed May 26, 2026).
Potter, Yujin, Wenbo Guo, Zhun Wang, Tianneng Shi, Hongwei Li, Andy Zhang, Patrick Gage Kelley, et al. 2025. "Frontier AI's Impact on the Cybersecurity Landscape." arXiv preprint arXiv:2504.05408.
Pratt, Jacob, and Albert Tanjaya. 2026. 2026 Transparency Report on Foundation Model Impacts: A Progress Report on Post-Deployment Governance Practices. Partnership on AI. https://partnershiponai.org/resource/2026-transparency-report-on-foundation-model-impacts (accessed on May 4, 2026).
Sela, Eyla. 2026. "A Single Operator, Two AI Platforms, Nine Government Agencies: The Full Technical Report." Gambit Security. https://gambit.security/blog-posts/a-single-operator-two-ai-platforms-nine-government-agencies-the-full-technical-report (accessed May 4, 2026).
Syed, Toqeer Ali, Mishal Ateeq Almutairi, and Mahmoud Abdel Moaty. 2025. "Toward Trustworthy Agentic AI: A Multimodal Framework for Preventing Prompt Injection Attacks." arXiv preprint arXiv:2512.23557.
Teun, van der Weij, Hofstätter Felix, Jaffe Ollie, Brown Samuel F., and Ward Francis Rhys. 2025. "AI Sandbagging: Language Models can Strategically Underperform on Evaluations." arXiv preprint arXiv:2406.07358.
Thornton, Scott. 2026. "Can Adversarial Code Comments Fool AI Security Reviewers: Large-Scale Empirical Study of Comment-Based Attacks and Defenses Against LLM Code Analysis." arXiv preprint arXiv:2602.16741.
TrustArc. 2025. California's AI Transparency Laws: How SB 942 and AB 2013 Will Reshape AI Data Practices. https://trustarc.com/resource/california-ai-transparency-laws-sb942-ab2013 (accessed May 4, 2026).
Vassilev, Apostol, Alina Oprea, Alie Fordyce, Hyrum Anderson, Xander Davies, and Maia Hamin. 2025. Adversarial Machine Learning: A Taxonomy and Terminology of Attacks and Mitigations. NIST AI 100-2e2025. Gaithersburg, MD: National Institute of Standards and Technology. https://doi.org/10.6028/NIST.AI.100-2e2025.
Wang, Hao, Quiyang Mang, Alvin Cheung, Koushik Sen, and Dawn Song. 2026. "How We Broke Top AI Agent Benchmarks: And What Comes Next." UC Berkeley Center for Responsible, Decentralized Intelligence (blog). April 2026. https://rdi.berkeley.edu/blog/trustworthy-benchmarks-cont/.
Wang, Zhun, Tianneng Shi, Jingxuan He, Matthew Cai, Jialin Zhang, and Dawn Song. 2025. "CyberGym: Evaluating AI Agents' Real-World Cybersecurity Capabilities at Scale." arXiv preprint arXiv:2506.02548.
Wang, Zhun, Nico Schiller, Hongwei Li, Srijiith Sesha Narayana, Milad Nasr, Nicholas Carlini, Xiangyu Qi, et al. 2026. "ExploitGym: Can AI Agents Turn Security Vulnerabilities into Real Attacks?" arXiv preprint arXiv:2605.11086.
Wei, Jason, Yi Tay, Rishi Bommasani, Colin Raffel, Barret Zoph, Sebastian Borgeaud, Dani Yogatama, et al. 2022. "Emergent Abilities of Large Language Models." arXiv preprint arXiv:2206.07682.
World Economic Forum. 2024. Strategic Cybersecurity Talent Framework. World Economic Forum.