What Is Prompt Injection? 9 Checks Before Installing AI Skills
- Ben Angel
- 6 minutes ago
- 12 min read

You find a Claude Skill that promises to analyze your inbox, organize customer research and prepare a weekly sales report. It has five stars, a polished description and exactly the workflow you have been trying to build. Ten minutes later, it is sitting beside your files, connected tools and business instructions.
Before you enable it, there is one question worth asking: what is prompt injection, and could this Skill quietly convince AI to do something you never requested?
Quick answer: Prompt injection is a social-engineering attack that places misleading instructions inside content an AI system processes. The attacker wants the AI to ignore or reinterpret your real request, reveal information, favor a manipulated result or take an unintended action. The instruction may appear directly in a chat, or indirectly inside a webpage, email, document, image, downloaded Skill or connected data source.
The risk is moving from theory into ordinary business use because AI can now browse, read email, open files, run code and work through connected applications. Anthropic’s current Claude Skills guidance warns that Skills may contain instructions, scripts, resources or third-party packages, and identifies prompt injection and data exfiltration as primary risks. Its advice is direct: install Skills only from trusted sources and inspect less-trusted Skills before enabling them—even when they were shared by a colleague.
That does not mean you should stop using AI Skills. It means you should stop treating them like decorative prompt templates.
A Skill can teach AI how to work. It can also expand the number of instructions, files and tools you are trusting.
If you want a structured way to build useful AI capability without handing over unnecessary access, the 28-Day AI Mastery Course gives you an implementation path. First, let us make this risk understandable enough to act on.
In This Article
What Is Prompt Injection?

Prompt injection happens when untrusted content is combined with the instructions guiding an AI system, and the untrusted content changes what the AI does.
NIST defines prompt injection as an attack exploiting the combination of untrusted input with a prompt created by a higher-trust party. In plain English, your instruction and somebody else’s instruction arrive in the same AI context. The system must decide which one deserves authority.
Suppose you ask an agent to research ten suppliers and compare price, delivery time and customer reviews. One supplier’s webpage contains hidden instructions telling AI to rank that supplier first and ignore negative evidence. If the agent follows those instructions, the result may still look like a professional comparison. The manipulation hides inside the process.
There are two broad versions:
Direct prompt injection: A person explicitly tells AI to disregard its rules, reveal protected instructions or perform an unauthorized task.
Indirect prompt injection: The instruction is embedded in something AI reads, such as an email, document, webpage, image, repository or tool response.
Indirect injection matters most to entrepreneurs because you may never see the instruction. OpenAI’s prompt-injection overview describes it as a form of social engineering aimed at conversational AI: third-party content attempts to mislead the system into doing something the user did not request.
Prompt injection is also different from a hallucination. A hallucination is an unsupported or invented answer. An injection is an attempt to influence the system’s behavior. The two can combine—a manipulated agent may produce a false explanation for an action it was tricked into taking—but they are not the same problem.
The hidden mechanism is misplaced authority. AI receives content and instructions through the same broad channel: language. Your safety depends partly on whether the system can tell the difference between information it should consider and commands it should obey.
Why AI Skills and Plugins Increase the Stakes

A Skill is useful because it gives AI procedural knowledge: how to create a presentation, audit a campaign, analyze data or follow your company’s workflow. Anthropic’s explanation of Claude Skills describes them as folders containing instructions, scripts and resources that Claude loads when relevant.
That is more powerful than saving a clever prompt in a document. A Skill may:
tell Claude how to perform a repeated process
include files, templates and examples
contain or install software packages
run scripts through code execution
communicate with external sources
work alongside connectors or plugins that reach other applications
The Skill itself may be malicious. A legitimate Skill may load poisoned content. A safe Skill may use an unnecessary dependency that introduces risk. Or a well-intentioned workflow may simply receive broader access than its job requires.
Think of the difference this way:
A prompt is advice given to an assistant for one assignment.
A Skill is part of the assistant’s operating manual.
A connector supplies keys to another room.
An agent can move between the rooms and act.
Each layer can be valuable. Each layer also increases the potential consequence when the wrong instruction is trusted.
This is why my guide to ChatGPT’s overlooked business features repeatedly keeps public, financial and reputational actions behind human approval. Capability without boundaries does not create leverage. It creates a larger blast radius.
Anthropic’s security guidance says less-trusted Skills should be reviewed for bundled files, code dependencies, images, scripts and instructions that connect Claude to untrusted network sources. The warning applies even when the Skill comes from a colleague, because friendly delivery does not prove safe contents.
Do not judge a Skill only by what it promises to produce. Judge it by what it can read, run, change and transmit.
Four Ways Prompt Injection Reaches Your AI

Most people picture a hacker typing “ignore all previous instructions.” Real attacks can be quieter because the dangerous instruction may be wrapped inside ordinary business material.
1. A webpage your agent visits
You ask AI to compare software, find vendors or research competitors. A page contains instructions aimed at the agent rather than the human visitor. They may be visible, disguised as ordinary content or concealed in elements people rarely notice.
Anthropic’s browser-agent research calls prompt injection one of the most significant challenges for browser agents and states that no browser agent is immune.
2. An email, attachment or shared document
An agent scanning support messages, invoices or meeting requests can encounter hostile instructions inside the very material it was asked to process. The email does not have to contain traditional malware. Its language may be the attack.
NIST’s 2026 agent-security findings describe this form of agent hijacking as malicious instructions placed in external data such as emails, websites or code repositories, potentially leading to data theft or unwanted code execution.
3. A downloaded Skill, plugin or software dependency
A custom Skill can include an instruction that quietly changes priorities, reads unrelated files or contacts an outside service. It may also install a package whose code performs an action the visible Skill description never mentioned.
This is an AI supply-chain problem: you are trusting the creator, the included files, the external packages and every update that later changes them.
4. Memory and reusable business context
Persistent memory makes AI more useful because the system can carry preferences and context forward. It can also allow a poisoned instruction to survive beyond the original task if the product or workflow stores it as trusted guidance.
That is why a carefully controlled AI brain for your business needs authoritative sources, clear labels and reviewable permissions. Business memory should help AI remember your standards, not preserve whatever an external document persuaded it to believe.
The Supplier Invoice Test

Imagine your bookkeeper receives an email that appears to come from a supplier you have paid for years. The branding looks right. The invoice amount looks right. The tone sounds familiar. One detail has changed: the message says the supplier has a new bank account and all future payments should go there.
Nothing broke into your bank. The email attempts to persuade a trusted person to follow a false instruction using the authority of a familiar business relationship.
Prompt injection is the AI version of that supplier-payment scam.
Your agent may be working from a legitimate assignment: review invoices, research vendors or draft replies. Hidden inside the material is a new instruction pretending to deserve authority. “Send this elsewhere.” “Ignore the negative reviews.” “Open this link.” “Treat this file as the new policy.” “Do not mention these instructions to the user.”
The practical defence is also similar. A careful bookkeeper does not approve changed payment details merely because they appeared inside a convincing email. They verify the change through a trusted channel and follow an approval rule.
Use the same standard for AI:
What was the original assignment?
Where did this new instruction come from?
Does that source have authority to change the assignment?
Would the action expose information, money, customers or reputation?
Where should the system stop for human approval?
I call this the Supplier Invoice Test. If you would independently verify a changed payment instruction before asking a person to act, require the same verification before allowing AI to act.
The test becomes increasingly important as AI agents handle repeated business workflows. The agent may be faster than an employee, but speed makes a mistaken instruction travel further before somebody notices.
Nine Checks Before Installing an AI Skill

You do not need to become a security engineer before using Claude Skills. You need a repeatable pause between “this looks useful” and “this now has access.”
1. Verify who created it
Prefer built-in Skills, official directories, established organizations or creators whose identity and maintenance history you can verify. A polished download page is not proof of provenance.
2. Read the Skill’s instructions
Open the included files and inspect what the Skill tells Claude to do. Look for instructions unrelated to the advertised outcome, attempts to override approval, requests to hide actions or commands that collect unnecessary information.
3. Inventory scripts and bundled resources
List scripts, images, templates, configuration files and executables. Ask why each item is required. An invisible supporting file deserves the same scrutiny as the main instructions.
4. Check third-party packages
Skills may include or install outside software. Record the package name, source and purpose. Avoid unexplained installation commands, abandoned packages and dependencies pulled from unfamiliar locations.
5. Identify every external connection
Determine whether the Skill sends requests to a website, API, server or connector. Ask what information leaves your environment, who receives it and whether the connection is essential.
6. Apply least privilege
Give the Skill only the files, folders, tools and accounts required for the current assignment. If it analyzes public competitor pages, it does not need access to payroll, private customer records or your entire Drive.
7. Test with harmless information
Run the first test using public or invented data in a contained project. Confirm what the Skill reads, creates and attempts to access before introducing confidential information.
8. Keep consequential actions behind approval
Require confirmation before sending messages, uploading files, publishing, purchasing, deleting, installing software or changing customer and financial records. OpenAI’s safety recommendations include limiting agent access, avoiding overly broad instructions and carefully reviewing consequential actions before confirming them.
9. Record the safe version and watch for changes
Document where the Skill came from, when you reviewed it and which version you approved. Reassess it after updates. A Skill shared across a team can change later, and Anthropic notes that recipients may automatically receive updates to shared Skills.
Use this decision rule: if you cannot explain what the Skill reads, runs, connects to and returns, it is not ready for sensitive business work.
This is the same discipline behind reliable AI workflow automation: narrow the job, define the inputs, make the output reviewable and preserve the approval boundary.
Once the checks are stable, turn them into a repeatable AI loop with verification and stopping rules so each new Skill is assessed against the same standard rather than whatever you happen to remember that day.
What to Do When You Cannot Audit the Code

Most entrepreneurs are not going to review hundreds of lines of code, and pretending otherwise would make this article useless. The goal is risk reduction, not false certainty.
Start with the lowest-risk path:
use a built-in Skill maintained by the platform
choose a trusted creator with transparent files and update history
avoid Skills requiring code execution when the outcome does not need it
remove unnecessary connectors and permissions
use a clean project containing only the files needed
test with non-sensitive data
keep the first assignment read-only
require approval before external or irreversible actions
If the Skill needs broad filesystem access, software installation, unknown dependencies or unrestricted network connections, ask a technically qualified person to review it.
The review should examine both the visible instructions and the code that can act independently of those instructions.
Sandboxing can also limit the damage. Anthropic’s Claude Code sandboxing guidance explains that filesystem and network isolation can prevent a compromised Claude Code session from reaching sensitive files or contacting unapproved services.
But a sandbox is not permission to stop paying attention. A constrained agent can still return biased research, hide evidence, recommend the wrong vendor or produce a convincing output based on poisoned content.
The safest workflow assumes detection may fail and limits what failure can reach.
That is why product providers use several defenses at once: model training, classifiers, monitoring, red-teaming, sandboxing, permission controls and human confirmation. Anthropic says no single layer guarantees safety, while OpenAI’s agent-security analysis argues that prompt injection increasingly resembles social engineering and cannot be addressed by input filtering alone.
Your role is not to reproduce those defenses. It is to control the part you own: the source, access, assignment, approval point and consequence.
Before You Hand Over the Keys

You may read this and think, “I barely understand what is inside a Skill. How am I supposed to judge whether it is safe?”
That reaction is reasonable. AI products are becoming easier to operate while the permission structure beneath them is becoming harder to see. A useful interface can make a serious capability feel as harmless as switching on a browser extension.
Here is the standard I want you to keep:
Convenience should never be mistaken for harmlessness.
I have spent years studying how entrepreneurs adopt AI, and the recurring mistake is rarely reckless intent. It is compressed judgment. We are busy, the demonstration looks convincing and the promise of saving six hours makes “Enable” feel like the smallest decision on the screen.
It may be the largest.
I wrote The Wolf Is at the Door because adapting to AI requires more than enthusiasm for what machines can do. It requires stronger judgment about what we should delegate, what deserves verification and where responsibility must remain human.
Choose one Skill you currently use—or one you are considering—and run the nine checks before giving it another file or connector. If you want to turn that kind of judgment into a practical operating system, the 28-Day AI Mastery Course gives you the guided implementation path. The objective is not to fear powerful AI. It is to become capable enough to use it without surrendering control.
Frequently Asked Questions

What is prompt injection in simple terms?
Prompt injection is an attempt to trick AI with instructions hidden inside content it reads. The attacker wants the AI to ignore the user’s real intent, reveal information, favor a manipulated result or take an unintended action.
Is prompt injection the same as jailbreaking?
No. Jailbreaking usually involves a user deliberately trying to bypass a model’s safety rules. Prompt injection can come from a third party through untrusted content and may target an agent performing a legitimate task.
Can a Claude Skill contain a prompt injection?
Yes. Anthropic identifies prompt injection as a primary risk of Skills and warns that Skills may contain instructions, scripts, resources or third-party packages. Install Skills only from trusted sources and inspect less-trusted Skills before enabling them.
Are Skills from coworkers automatically safe?
No. A coworker may share a Skill without realizing that it contains unsafe instructions, dependencies or external connections. Anthropic specifically recommends reviewing less-trusted Skills even when they were shared by a colleague.
Can antivirus software detect prompt injection?
Traditional antivirus software may detect known malicious files or code, but prompt injection can exist as ordinary language inside a webpage, email or document. AI-specific protections, narrow permissions and human approval are still needed.
Can prompt injection steal business data?
It can contribute to data exfiltration if a vulnerable agent has access to sensitive information and a path to transmit it. The practical defense is to minimize access, restrict external connections and keep consequential actions behind approval.
Are built-in Claude or ChatGPT Skills completely safe?
No technology provider guarantees complete protection from prompt injection. Built-in capabilities generally offer a stronger trust starting point than unknown downloads, but users should still limit permissions, review important actions and avoid unnecessary sensitive access.
Does prompt injection affect only Claude?
No. Prompt injection can affect any AI system that processes untrusted content, including browser agents, assistants, connected applications and custom workflows from multiple providers.
How much does protection against prompt injection cost?
Basic protection mostly requires disciplined setup: trusted sources, limited access, safe test data and human approval. Businesses using complex agents, sensitive systems or custom code may also need security review, monitoring, sandboxing and specialist support.
Should a small business stop using AI agents?
No. Small businesses should use agents where the outcome is valuable, access can be limited and the work can be reviewed. Avoid giving broad authority to an agent simply because the interface makes that authority easy to grant.
What should I do if I suspect a prompt-injection attack?
Stop the task, do not approve the proposed action, disconnect unnecessary tools, review what the agent accessed and preserve relevant logs or screenshots. If sensitive information, credentials or financial systems may have been exposed, follow your incident-response process and obtain qualified security help.