AI is moving beyond the chatbot era. Instead of simply answering a question, modern AI agents can browse websites, read documents, access applications, call APIs, write code, move information between services and complete multi-step tasks with limited human intervention. That shift is creating a new security problem: the biggest risk may no longer be what an AI system says, but what it is allowed to do.

Security researchers, technology companies and standards organisations are increasingly focusing on risks such as prompt injection, agent hijacking, excessive permissions, unsafe tool usage and compromised third-party connections. NIST has specifically identified agent hijacking through indirect prompt injection as an important security concern, while OWASP has introduced dedicated guidance for securing agentic applications.

The important change is simple: when an AI has access to real systems, a mistake or manipulated instruction can potentially become an action. An agent that only produces a wrong answer may waste someone's time. An agent connected to email, cloud storage, financial systems or developer tools could potentially cause a much larger problem.

Introduction: AI Is Learning to Take the Next Step

For years, the basic interaction with AI was straightforward. A user asked a question, the model generated an answer and the human decided what to do next. Agentic AI changes that relationship. The user can provide a goal and the system may determine the steps required to achieve it, select tools, retrieve information and execute actions along the way.

That makes AI considerably more useful, but it also changes the security boundary. An AI agent may process information that comes from sources the developer does not control, including emails, websites, documents, repositories and third-party applications. NIST's research on agent security highlights exactly this problem: malicious instructions can be hidden inside external data and influence an agent's behaviour, potentially causing it to expose information or perform unintended actions.

The result is a new question for technology teams. Instead of asking only whether an AI model is accurate, organisations increasingly need to ask whether the agent can be trusted with the permissions, tools and data required to complete its job.

Why Agentic AI Changes the Security Equation

The fundamental difference is agency. A conventional AI assistant may recommend that a user send an email. An agent could potentially draft it, access the relevant contact and send it through an integrated service. A coding assistant may go beyond suggesting code and modify files, execute commands or interact with development tools.

This is where the concept of excessive agency becomes important. OWASP describes the problem as giving an AI system excessive functionality, permissions or autonomy, allowing unexpected or manipulated model outputs to result in damaging actions. The danger can arise from hallucinations, prompt injection, compromised extensions or poorly designed workflows.

The issue is therefore not simply that AI models can make mistakes. Software has always contained bugs. The difference is that an AI agent can dynamically decide how to use the tools it has been given, sometimes across several steps. A small failure at the beginning of a workflow can therefore propagate through the rest of the process.

Understanding the Attack: What Is Prompt Injection?

One of the most important concepts in agent security is prompt injection. In simple terms, it happens when malicious instructions are placed inside information that an AI system is expected to process.

Imagine an employee asks an AI agent to review their inbox and summarise important emails. One of those emails contains hidden instructions telling the agent to ignore its original task and forward sensitive information somewhere else. A vulnerable agent may interpret those instructions as part of the task rather than treating them as untrusted content.

This is known as indirect prompt injection because the attacker does not necessarily communicate directly with the AI system. Instead, malicious instructions are embedded in a webpage, document, email or another data source that the agent later encounters. Microsoft describes potential consequences ranging from data exfiltration to unintended actions performed using the user's credentials.

NIST has also studied agent hijacking, describing scenarios where malicious instructions inserted into external data can redirect an AI agent toward harmful behaviour. This makes every external information source part of the agent's potential attack surface.

The Real Problem Is Not Just the Model

It is tempting to treat agent security as a model problem: make the model smarter and the security issue will disappear. The evidence suggests that this is not enough.

An agent is an entire system consisting of the model, instructions, memory, tools, connectors, authentication mechanisms, data sources and surrounding application infrastructure. A stronger model can reduce certain failures, but it does not automatically make every connected tool safe.

Anthropic, for example, has described prompt injection as an ongoing challenge and recommends thinking about security across multiple layers rather than depending on one defence. Its research also emphasises carefully controlling the tools and data made available to an agent.

This leads to an important engineering principle: do not give an AI agent more authority than its task requires.

If an agent only needs to read a document, it should not automatically receive permission to delete documents. If it needs to schedule meetings, it may not need unrestricted access to every calendar operation. If it needs to retrieve customer information, it should not automatically have permission to modify customer records.

The New Security Stack for AI Agents

Securing an agent requires more than placing a firewall around an AI model. The protection has to follow the agent throughout its workflow.

At the identity layer, organisations need strong authentication and tightly scoped permissions. At the data layer, sensitive information should be separated according to who and what can access it. At the tool layer, every connector and API should expose only the functions the agent genuinely needs.

The execution layer also needs monitoring. Organisations should be able to determine what an agent accessed, which tools it called, what decisions it made and what actions followed. Microsoft recommends treating autonomous agents as a distinct workload and highlights the importance of governance, authorization and monitoring as agent deployments expand.

This is where human approval remains valuable. High-impact operations such as deleting data, transferring money, changing production infrastructure or sending sensitive communications should not necessarily be completely autonomous. The most practical architecture may be one where AI handles routine decisions while humans remain in the loop for irreversible or high-risk actions.

Why MCP Makes the Conversation More Important

The growth of protocols such as the Model Context Protocol (MCP) is making it easier for AI applications to connect with external tools and information sources. That is useful because agents become much more capable when they can interact with real systems.

But every new connection can also introduce another trust boundary. Microsoft notes that poorly governed MCP implementations can create risks involving data exfiltration, prompt injection and unvetted services, making granular access control increasingly important.

The challenge is therefore not to stop agents from using tools. That would remove much of their value. The challenge is to build a permission system in which agents can use tools without receiving unrestricted authority over everything those tools can do.

What This Means for Developers

For developers, the rise of agentic AI changes the way applications need to be designed. Traditional application security practices still matter, but they need to be combined with controls specifically designed around probabilistic AI behaviour.

Microsoft recommends treating model outputs as untrusted and carefully validating data flows, tool configurations and external content. Its guidance specifically warns about indirect prompt injection through retrieved documents and other context sources.

A practical development workflow should therefore include threat modelling before deployment, least-privilege permissions, strict tool definitions, input and output validation, logging, approval gates for sensitive actions and continuous testing against adversarial inputs.

Developers should also test the complete agent rather than evaluating only the underlying model. An impressive benchmark score does not tell you whether an agent can safely operate an email account, production database or cloud environment.

Real-World Impact: From Chatbots to Digital Employees

The implications extend far beyond AI companies. Consider a software company where an AI agent can inspect GitHub issues, modify code, run tests and create pull requests. The productivity gains could be significant, but so is the potential impact of a compromised workflow.

The same principle applies to customer support, finance, healthcare, education and enterprise operations. An AI agent connected to internal systems can potentially become a powerful interface to organisational data and services. That makes identity, access control and monitoring central parts of the AI architecture rather than secondary security features.

The business question is therefore changing from “Can we automate this task?” to “What level of autonomy can we safely give the system?”

What Companies Should Do Now

Companies adopting AI agents do not necessarily need to wait for perfect security technology. They can start by mapping exactly what each agent can access and what actions it can perform.

The next step is to reduce unnecessary permissions, separate sensitive systems, require confirmation for high-impact operations and maintain detailed activity logs. Security teams should also conduct adversarial testing in which researchers deliberately attempt to manipulate agents through webpages, documents, emails and tool outputs.

NIST's 2026 analysis of industry responses on AI-agent security found broad agreement that agents introduce novel security threats and that traditional cybersecurity practices will need to be adapted for this new environment.

For developers building agents today, this creates a useful checklist: minimum permissions, trusted tools, isolated execution, continuous monitoring, human approval where necessary and regular red-team testing.

The Outlook: The Future May Belong to Controlled Autonomy

Agentic AI is unlikely to disappear because of these security challenges. In fact, the ability to delegate multi-step work is one of the strongest reasons organisations are interested in agents in the first place.

The more realistic future is therefore not fully autonomous AI operating without restrictions. It is controlled autonomy: systems that can act independently within clearly defined boundaries and know when a decision needs human approval.

The industry is already moving in that direction. OWASP has created a dedicated Top 10 for agentic applications, NIST is researching agent hijacking and security evaluation, and major technology companies are developing their own defensive approaches.

The central challenge will be designing agents that are powerful enough to be useful without becoming powerful enough to create unacceptable consequences when something goes wrong.

Knowledge Corner

Agentic AI: An AI system that can plan and execute multiple steps toward a goal, often using external tools or applications.

Prompt Injection: An attack in which malicious instructions attempt to influence an AI model's behaviour.

Indirect Prompt Injection: A prompt-injection attack delivered through external content such as webpages, documents, emails or retrieved data.

Excessive Agency: A situation where an AI system has more functionality, permissions or autonomy than it needs, increasing the potential impact of mistakes or manipulation.

MCP: An open protocol designed to help AI applications connect with external data sources and tools. Its flexibility also makes access control and governance important.

What Should Developers Learn From This?

The biggest lesson is that building an AI agent is no longer just about choosing a good model and writing a clever prompt. The difficult part is designing the environment around the model.

For students and developers entering AI engineering, security should become part of the core skill set rather than something added after deployment. Understanding authentication, authorization, API security, sandboxing, logging, threat modelling, prompt injection and evaluation can make an AI project substantially more production-ready.

The future AI engineer may therefore look less like someone who simply integrates an LLM API and more like a systems engineer who understands models, software, infrastructure and security together.

Expert FAQs

Is agentic AI more dangerous than a normal chatbot?

Not inherently, but the potential impact of a failure can be much greater when an agent has permission to perform real actions. A chatbot may generate an incorrect answer, while an agent could potentially use connected tools to act on that incorrect or manipulated output.

Can prompt injection be completely prevented?

Current research does not support treating prompt injection as a solved problem. Anthropic notes that increasingly capable agents still face prompt-injection risks and that layered defences are necessary.

Should AI agents always have human approval?

Not necessarily. Requiring approval for every low-risk action can eliminate much of the benefit of automation. A better approach is to classify actions according to risk and introduce human approval for sensitive, irreversible or high-impact operations.

What is the most important security principle for an AI agent?

Least privilege is one of the most practical principles: give an agent only the permissions and tools it actually needs. Reducing unnecessary authority limits the potential damage if the model makes a mistake or becomes compromised.

Is traditional cybersecurity still useful for agentic AI?

Absolutely. Authentication, authorization, encryption, monitoring, secure software development and access control remain important. The difference is that these practices need to be adapted to systems where decisions and actions can be generated dynamically by AI. NIST's 2026 analysis similarly found that established cybersecurity practices remain relevant but require adaptation for AI agents.

Final Takeaway

The next major AI security battle may not be about stopping a model from producing a bad answer. It may be about stopping an AI agent from turning a bad instruction into a real-world action.

That is why the rise of agentic AI should be viewed as an infrastructure and security shift, not simply another upgrade in chatbot capability. As agents gain access to email, code, cloud platforms, enterprise databases and financial systems, the question of who controls the agent, what it can access and when it is allowed to act becomes just as important as the intelligence of the model itself.

The winners of the agentic era will not simply be the organisations with the most capable AI. They will also be the ones that can make that AI powerful without giving it unnecessary power.