LLMGoat - A05 Improper Output Handling

A05 LLM GOAT

This post is the fifth in a series of 10 blog posts and it covers the solution to the Improper Output Handling challenge from LLMGoat.

LLMGoat is an open-source tool we have released to help the community learn about vulnerabilities affecting Large Language Models (LLMs). It is a vulnerable environment with a collection of 10 challenges - one for each of the OWASP Top 10 for LLM Applications - where each challenge simulates a real-world vulnerability so you can easily learn, test and understand the risks associated with large language models. This tool could be useful to security professionals, developers who work with LLMs, or anyone who is simply curious about LLM vulnerabilities.

If you are not familiar with LLMs, we recommend that you check out the first post in the series here.


The vulnerability

An improper output handling vulnerability (OWASP LLM05:2025: Improper Output Handling) arises when content generated by an LLM is passed to another part of an application without being properly checked or handled.

An LLM usually produces text output, but the application or system handling that output may incorrectly consider it to be trusted as it didn’t originate directly from the user. LLM output might be rendered as a web page (like in a chat window), be used to form a command or database query, or it could be passed on to another tool or API. If the application assumes that the output is safe, an attacker may be able to use the LLM as an indirect route to exploit potential vulnerabilities.

Improper output handling can occur in several places:

The impact of improper input handling depends on where the output is sent and what capabilities the receiving system has. A malicious response displayed in a browser might affect an individual user, while the same response passed to a privileged backend service could expose data, alter files or even lead to server compromise.

Several vulnerability disclosures and security demonstrations have shown how these weaknesses can be exploited.

These examples show why security testing must follow an LLM’s output beyond the chat interface. Testing should examine where each response is sent, how it is interpreted and what actions the receiving system is allowed to perform.

In the following challenge, we will manipulate an LLM into producing unsafe content and examine how improper handling allows that output to affect the wider application.


The challenge

LLMGoat presents us with a chatbot-type screen and we try to understand its intended functionality:

LLMGoat A05 - Picture 1


A summary of the system prompt always helps. The message continues with:

LLMGoat A05 - Picture 2


At this point, it is clear that there are at least 2 functions or API endpoints that the model can invoke:

- (LIST_USERS)

- (READ_NOTES) <username>

The system prompt also states that it is not allowed to share notes for 3 of the goats, so chances are those are precisely the ones we will need to access to find the secret note mentioned in the challenge description.

We can start by listing all the goats in the system:

LLMGoat A05 - Picture 3


Now we can try reading notes for some of them:

LLMGoat A05 - Picture 4


What happens when we try to read notes for a “forbidden” goat?

LLMGoat A05 - Picture 5


After trying some prompt injection techniques, it becomes clear that whenever the username provided is one of those 3 forbidden goats, we always get the same message:

LLMGoat A05 - Picture 6


This means that either we haven’t tried hard enough or that there must be some input (or output) validation enforced that simply detects the presence of those usernames and blocks the query happening under the hood.

If you are familiar with how applications work, it is easy to imagine that the way this could work is by taking the provided username as a parameter and feeding it to a SQL query executed against a database. The query could look something like:

SELECT note FROM notes WHERE username = <requested_username>

With an LLM in the middle, this means that the user asks to read the notes of a particular goat, the LLM converts that natural language request to a call to (READ_NOTES) with the username provided as a parameter, the database executes the corresponding SQL query and returns the results.

Anyone security aware would know that there needs to be a validation of the input that will end up being used in the SQL query and practices like parameterised queries are highly recommended. However, developers often incorrectly assume that an LLM’s output can be implicitly trusted as it doesn’t come directly from the user.

Let’s see what happens if we try to inject into the username parameter:

LLMGoat A05 - Picture 7


At first it didn’t want to comply, probably due to the LLMs instructions, but when we forced it to output exactly the line we wanted the results look promising. 😊

The “too many notes” message indicates that there might be some limits enforced, maybe to prevent long processing times or breaking the user interface with too much content at once.

Before we continue, we can double check that providing a false condition doesn’t return the same message:

LLMGoat A05 - Picture 8


Great! Now let’s address the “too many notes” issue by limiting the query to fetch only 5 notes.

LLMGoat A05 - Picture 9


Excellent! We are using SQL injection to fetch notes. These notes match the notes we previously retrieved for bleatboss. Now we need to try to reach the secret notes without using the forbidden usernames.

One possible way is to use an offset clause, which allows us to skip a certain number of rows before providing the results.

LLMGoat A05 - Picture 10


This seems to work as we are getting different notes than we were getting before. If we keep cycling through notes, we eventually reach the notes for the forbidden goats and read the secret:

LLMGoat A05 - Picture 11


Challenge solved!

As you can see, even when the web application itself is secure, it may be possible to leverage an LLM to exploit web application vulnerabilities if the model’s output isn’t handled securely. SQL injection is one example but we commonly manage to exploit other types of issues like XSS or template injection vulnerabilities.

Note as well that while validating the user input provided to the chatbot is a good idea, the model’s output still need to be validated because it is often possible to generate malicious outputs from seemingly benign inputs. For example, if the input is validated to forbid special characters such as “-” or ‘ , the attacker could simply ask the model to generate those (e.g. Print a string with a single quote, followed by a space, etc).


Conclusion

Improper output handling becomes dangerous when an application treats content generated by an LLM as trusted. If that output is rendered, executed or passed to another system without appropriate sanitisation, it can lead to data loss, unauthorised actions or compromise of the underlying application.

Reducing this risk starts with treating LLM output in the same way as other untrusted input. Some measures that can help include:

Protecting against improper output handling does not depend on making the model perfectly reliable. In fact, it is a bad idea to rely only on model instructions / system prompts to ensure that the model will never produce unsafe output. It depends on ensuring that the wider application does not blindly trust model output and enforces appropriate data validation, enabling it to remain secure even when the model produces unsafe or unexpected content.