Artificial intelligence arrives in the cybersecurity domain at a moment in which defenders are already contending with adversaries that are faster, more adaptive, and more resource efficient than they are. The preceding sections have argued that recent advances in generative AI, reasoning, and agentic systems will, in the short term, deepen this asymmetry, while also enabling the conditions for a more favorable defensive posture over longer time horizons.
AI represents a turning point for cybersecurity, with near-term risks and long-term potential. Frontier AI systems are rapidly expanding what is possible in cybersecurity—for both attackers and defenders. As discussed earlier, near-term AI advances are likely to favor attackers by reducing the time, expertise, and operational effort required for key cyber activities.
This imbalance is compounded by the rapid pace of capability development and diffusion. Even if leading providers restrict access to their most powerful models, comparable capabilities are likely to emerge through both open-weight models and globally distributed development efforts within months, not years. At the same time, AI introduces new attack surfaces, including vulnerabilities associated with AI system deployment, as well as amplified social engineering threats enabled by highly personalized and convincing synthetic content.
The central challenge is operating under uncertainty while capabilities outpace measurement. A major obstacle to effective response is the difficulty of measurement. AI-driven cyber capabilities are advancing faster than the ability to evaluate them, especially in the context of security posture. There are no widely accepted benchmarks for assessing offensive or defensive performance, nor for tracking how quickly those capabilities are improving over time. More fundamentally, the full impact of generative AI systems is inherently unpredictable and limiting the ability of policy makers and practitioners to anticipate failure modes or bound risk.
Addressing this gap will require coordinated efforts across multiple communities. Formal partnerships among frontier AI developers could support the co-development of shared capability assessments, enabling more consistent evaluation of model performance in cybersecurity-relevant tasks. Public–private collaborations—bringing together academia, national laboratories, NIST, and other trusted research institutions—can play a central role in developing independent testing methodologies, benchmarking frameworks, and evaluation infrastructure. Expanding controlled access to leading models for qualified researchers, balanced against the incremental risks of wider exposure, and paired with targeted research funding, can further accelerate the development of evaluation approaches that assess not only capability, but also robustness and behavior under adversarial conditions.
This lack of measurement complicates decision-making and reinforces longstanding misalignments between incentives and security outcomes. At the same time, the growing inevitability of AI-enabled attacks strengthens the case for investments in such practices as red teaming, adversarial testing, and continuous evaluation. Improved measurement and shared assessment frameworks would help organizations calibrate responses more effectively while reducing the risks of both underreaction and misaligned intervention.
Cybersecurity must improve rapidly, while some interventions may buy time. Even amid this uncertainty, it is clear that the baseline level of cybersecurity across society will need to rise. Organizations will need to adopt stronger security practices, improve software quality, invest in more resilient architectures, and respond more quickly to emerging threats. AI can support this transition by enabling faster detection, improved communication among