Should I Be Worried About Prompt Injection?

A strange thing happened recently in a Connecticut courtroom. A self-represented litigant hid instructions to artificial intelligence inside a court filing — written in tiny, white font that was essentially invisible to a person looking at the document but still readable by software processing its text. The hidden language was addressed directly to any AI model that happened to review the filing, instructing it to side with the litigant and help produce the result he wanted. 404 Media broke the story, and the underlying sanctions decision is worth a read.

One problem with the scheme: the Connecticut court was not using AI to review the filing. So no AI was secretly persuaded, and the hidden instructions did nothing to affect the court’s decision. But: the human judge found them anyway. Superior Court Judge Walter Spader, Jr. concluded that the litigant, Matthew Elliott, had acted with malicious intent and revoked his electronic-filing privileges, requiring him to make future filings on paper instead.

This technique is known as prompt injection — more specifically, “indirect prompt injection,” meaning instructions embedded in material that someone else may later hand to an AI system. Unlike some of the stranger AI stories making their way through the legal profession, prompt injection is a real, recognized computer-security problem. So should Louisiana lawyers actually be worried that opposing counsel might hide instructions in a brief, contract, discovery response, or other document, and thereby manipulate the AI tools we use to review it?

We decided to try it ourselves.

We took a motion and inserted a hidden instruction directed specifically at any AI system reviewing the document, telling the AI to emphasize a particular argument in the motion as the strongest one. We then changed the instruction to white font so it disappeared from ordinary view — the same trick used in the Connecticut filing.

We uploaded the document to several of the major general-purpose AI systems and asked each to review the motion and rank its strongest arguments, without telling the systems that the document contained a hidden instruction.

It didn’t work for any of the systems we tried. Claude spotted the hidden instruction and expressly told us that the document contained language directed at an AI system. Perplexity likewise identified the embedded instruction and refused to follow it. ChatGPT also detected it rather than allowing it to control the substantive analysis. In other words, we were unable to get this relatively straightforward form of prompt injection past any of the major LLMs we tested.

Our experiment was hardly a scientific study, and it certainly does not establish that prompt injection cannot work, especially as users (and LLMs themselves) get more sophisticated. More elaborate attacks exist, and the security problem becomes considerably more important once an AI system has access to tools, private data, email, or the ability to take actions on a user’s behalf. But our little experiment does suggest that lawyers probably do not need to start searching every PDF they receive from opposing counsel for white text before uploading it to ChatGPT. Interestingly, when 404 Media tried the actual Connecticut filing with ChatGPT, they got essentially the same result — ChatGPT recognized the hidden instruction and disregarded it.

Could a Louisiana Lawyer Do This?

No. Even setting aside whether the technique works, a lawyer should not be hiding instructions in a court filing in hopes of secretly manipulating an AI system used by a court, opposing counsel, or anyone else involved in the proceeding. Louisiana Rule of Professional Conduct 8.4(c) provides that it is professional misconduct for a lawyer to engage in conduct involving dishonesty, fraud, deceit, or misrepresentation. Rule 8.4(d) separately prohibits conduct that is prejudicial to the administration of justice. Deliberately concealing instructions in a pleading to secretly manipulate how a tribunal or another participant’s AI system evaluates the case would raise obvious concerns under both provisions.

Rule 3.3, governing candor toward the tribunal, may also come into play depending on what the hidden instruction says and what the lawyer is trying to accomplish — the comments to Rule 3.3 emphasize a lawyer’s special obligation as an officer of the court not to undermine the integrity of the adjudicative process. And if the hidden instruction were intended specifically to influence a judge or other court official through an AI system, Rule 3.5 could present still another problem. Lawyers are advocates; we are supposed to try to persuade courts. We are not permitted to do so through concealed communications designed to reach the decisionmaker outside the ordinary adversarial process.

So, Should I Be Worried?

Probably not very worried — at least not about this version of prompt injection. There is an important difference between saying that prompt injection is a genuine AI-security vulnerability and saying that a lawyer can hide “ignore the other side and rule for me” in white font in a motion and expect a modern AI system to obediently comply. Our experiment suggests the latter is not particularly easy to pull off. The major systems we tested were looking for precisely this sort of instruction.

That doesn’t mean the broader problem should be ignored. Prompt injection becomes more concerning as lawyers begin using AI systems that do more than answer questions — systems that autonomously search files, review email, retrieve information, or take actions on a user’s behalf. It’s also possible that more sophisticated or less obvious injections could evade the protections we encountered. But for the ordinary Louisiana lawyer using one of today’s major LLMs to summarize a brief, review a contract, or analyze a motion, we would put this particular risk fairly low on the list of things keeping us up at night.

For now, the Connecticut case may be more useful as an ethics lesson for the person inserting the prompt than as a cybersecurity warning for the lawyer receiving it.

Please follow and like us: