The rapid advancement of AI in cybersecurity presents a fundamental challenge: the ability to measure and evaluate risk has not kept pace with the development and deployment of new capabilities. While the current moment appears transformative, and in some respects unprecedented, it also reflects a familiar pattern—capability is relatively easy to demonstrate and quantify particularly for offensive capabilities, which can often be directly demonstrated and benchmarked, while defensive effectiveness and overall security posture are inherently more difficult to evaluate and quantify. Risk and security have always been difficult to measure in adversarial environments because vulnerabilities may be unknown and attacker behavior cannot be fully predicted or detected. However, the probabilistic and rapidly evolving nature of generative AI systems compounds these longstanding challenges and complicates efforts to assess the implications of AI for cybersecurity and to design effective responses.
A central difficulty in assessing AI-enabled cybersecurity risk is the absence of behavioral guarantees in generative AI systems. These systems are fundamentally approximation algorithms implemented as nonlinear estimators, and as such they lack strict performance bounds. Even experts are frequently unable to explain certain aspects of model performance in a formal sense, particularly in novel or adversarial contexts.
While useful approximations of performance can be developed for expected inputs, these characterizations are less reliable when systems encounter edge cases or intentionally adversarial inputs (Mohsin et al. 2025). This introduces uncertainty into systems that are increasingly being deployed in high-stakes environments.
At the same time, users—particularly nonexperts—may overestimate the reliability and robustness of AI outputs because system limitations are not always transparent or well understood. This gap between perceived and actual system behavior can complicate risk assessment and lead organizations to deploy or rely on AI-enabled systems in ways that introduces additional security risk.
Although benchmarks and evaluation frameworks for AI systems are rapidly being developed (AI Security Institute n.d.; NIST n.d.; Pratt and Tanjaya 2026), most focus on capability—whether a system performs according to specification on a given task. In cybersecurity, however, capability does not map clearly onto security outcomes. Security is defined not only by expected performance, but by resilience to failure, misuse, and adversarial behavior, as well as system behavior under deliberately manipulated conditions (Z. Wang et al. 2025; H. Wang et al. 2026). These properties are inherently more difficult to evaluate because assessments often depend on organizational context, and experts may reasonably disagree about whether a system's behavior is desirable or secure.
Frameworks can help structure analysis by modeling specific scenarios, but they are necessarily limited in scope (NIST 2023). Generative AI systems, by contrast, are designed to operate across a wide range of contexts, making it difficult for any single framework to capture the full range of possible behaviors and associated risks (NIST 2024). This limitation is compounded by the pace of capability development, which is beginning to outstrip existing evaluation frameworks, as rapidly evolving AI systems exceed the scope, assumptions, or difficulty levels embedded in current benchmarks (Otto 2026; Z. Wang et al. 2026). Other factors that confound assessment include answers that leak into training data and capability concealment, whereby an AI system intentionally underperforms on an
evaluation to obscure its true capabilities (Teun et al., 2025). As a result, there remains a gap between what can be measured and what must be managed.
Risks associated with AI-enabled cybersecurity fall into three categories: (1) risk associated with AI systems, (2) risk associated with human users, and (3) risks arising from human–AI interaction. These categories highlight an important limitation of current benchmarks and evaluation frameworks, as their limited coverage does not fully capture risks that emerge across these dimensions, especially involving human behavior and system interaction over time.
Key sources of risk include training, data, and system usage. Vulnerabilities can be introduced through both malice and oversight; from a security perspective, the intent behind a vulnerability is less important than its existence. For example, a model may fail to perform reliably because of insufficient exposure to certain data distributions during training or may produce erroneous outputs because of poorly aligned datasets. While not all such risks can be fully enumerated, many can be tested and evaluated extensively prior to deployment, particularly in controlled environments (OWASP Gen AI Security Project n.d.; Vassilev et al. 2025). This testing is particularly important in contexts sensitive to low-probability events, such as national security applications and comparable edge cases in other domains. In contrast, achieving meaningful testing coverage for adversarial inputs is significantly more difficult because of the enormous state space, including all possible system configurations, available to attackers.
A distinct and increasingly important category of risk arises from agentic AI systems and their interactions with other AI systems. Unlike traditional input–output models, agentic systems can access tools, retrieve and process data, and execute multi-step tasks over extended time horizons with limited human oversight. These systems are being granted increasing levels of agency and autonomy—and correspondingly introduce qualitatively different risks. Errors or misaligned decisions may compound over time in poorly quantifiable ways, and interconnected systems may exhibit emergent behaviors that are difficult to predict or control. In such settings, risk arises not only from individual system failures, but also from the dynamics of interacting agents operating in shared environments. At the same time, the increasing accessibility of advanced AI capabilities—through open-weight models, fine-tuning, and agentic frameworks—may democratize access to sophisticated cyber capabilities, enabling a broader range of actors to conduct complex operations. These trends suggest an expansion from isolated system vulnerabilities toward more systemic risks arising from autonomy, interaction, and scale, and highlight the importance of continued research into defining and implementing effective guardrails for increasingly autonomous systems.
Drawing on decades of cybersecurity experience, humans remain a significant vulnerability. Many major security incidents can be traced to human error, negligence, or susceptibility to manipulation. Generative AI systems amplify these risks by enabling more sophisticated and targeted social engineering attacks, including highly personalized phishing campaigns that can be generated and deployed at scale with progressively less effort and investment (see section on "Social Engineering and Synthetic Identity" below). Measurement of human-associated risk is further complicated by the difficulty of attribution and the variability of human behavior across contexts.
Human–AI interaction introduces both opportunities and new forms of risk. Research on human–AI teaming suggests that combined systems can outperform either humans or machines alone when properly designed, particularly when tasks are structured to leverage complementary strengths such as human judgement and machine-scale analysis. As a result, testing the performance of human–AI teams—in contexts ranging from coding to diagnosis to teaching—offers the potential to identify areas of weakness and vulnerability. However, these assessments will need to evolve alongside the integration of generative models, given the fluidity of the partnerships involved.
At the same time, these systems introduce new failure modes. Overreliance on AI systems may reduce human vigilance, increase automation bias, and lead users to defer to AI outputs without sufficient scrutiny, particularly when those outputs appear confident or authoritative. A further dimension of risk concerns the longitudinal impact of sustained reliance on AI on the development of human expertise. There is a meaningful difference between the use of AI by an expert—such as an experienced developer using generative tools to augment their workflow—and its use by a novice. The reduction of "learning by doing" may have long-term implications for the technical and analytical capabilities of the workforce, particularly if foundational skills are not developed or maintained. Over time, this could erode the human capacity required to understand, validate, and respond to complex cybersecurity challenges, especially in situations where AI systems fail or behave unpredictably.
Whether all risk associated with the introduction of generative AI can be measured is, to a significant degree, less important than whether the risks that can be measured are in fact measured. It is generally impossible to quantify all risk in all situations, but that should not prevent the systematic evaluation and mitigation of risks that are well defined.
At present, one of the most significant challenges is the misalignment of incentive structures with security outcomes. This problem predates generative AI—organizations have long prioritized the deployment of new capabilities over the rigorous evaluation of associated risks, particularly when measurement is difficult or costly. However, the rapid advancement and diffusion of AI capabilities make this misalignment more acute.
The pressure to develop increasingly capable AI systems, combined with the difficulty of assessing their security implications, reinforces the gap between capability and security and complicates efforts to improve overall resilience. At the same time, the growing inevitability of AI-enabled attacks strengthens the case for investing in practices such as red teaming and adversarial testing, providing a clearer justification for the costs associated with proactively identifying and mitigating vulnerabilities before they are exploited. AI-enabled development pipelines also present new opportunities to include security concerns from the early stages of development, as discussed further in the following section.
Addressing these incentive and measurement challenges will require coordinated efforts across sectors. Public–private partnerships can play a central role in improving measurement, sharing risk-relevant information, and developing common evaluation frameworks. Academic institutions, government agencies, and nonprofit organizations—which can contribute independent evaluation, long-term research perspectives, and public-interest considerations—are particularly well positioned to contribute to independent testing, benchmarking, and the development of methodologies for assessing AI-enabled cybersecurity risks. Industry participation is equally critical, given its access to systems, data, deployment environments, operational expertise, and incentives to
develop solutions that are deployable, effective, and scalable in practice (NASEM 2025). At the same time, the pace of AI development and the likely increase in AI-enabled attacks will require coordination mechanisms that support rapid sharing of threat intelligence and defensive practices. Shared frameworks, standards, and regulatory approaches may provide important benefits, but they must be designed carefully to avoid unnecessary rigidity, implementation burdens, or delays that could impede effective responses in a rapidly evolving technological environment.
In the absence of robust measurement, decision makers must operate under conditions of uncertainty. This observation underscores the importance of adopting approaches that are resilient to unknown risks, including conservative deployment strategies in high-consequence environments, sustained investment in evaluation infrastructure, and continued support for foundational research to better understand system behavior.