- Getting started
- UiPath Conversational Agents
- UiPath Agents in Studio Web
- UiPath Coded agents
- Build with Coding Agents
Out-of-the-box guardrails
Enable predefined guardrails for PII, prompt injection, harmful content, and IP leakage, or configure LLM as Judge to evaluate custom natural-language policies, then set actions to log, block, or escalate.
Out-of-the-box guardrails are predefined, ready-to-use safeguards that you can enable for your agents without any custom configuration or coding. They provide immediate protection against common risks such as sensitive data exposure and prompt injection attacks, helping you build secure and trustworthy agents faster.
The guardrails on this page run on UiPath-managed safety services. If an administrator has configured your organization's subscription to a supported safety vendor in the AI Trust Layer, its guardrails appear as additional options in the Choose guardrail panel, grouped under the folder they were created in and marked BYO, with the connector listed as their provider. You configure them the same way as the guardrails on this page.
For details, including which vendors are supported, refer to Configuring guardrails in the Automation Cloud admin guide.
Prerequisites
Out-of-the-box guardrails are available on the following licensing plans:
- Flex Pricing plan: Enterprise – Standard and Advanced tiers.
- Unified Pricing plan: Standard, Enterprise, App Test Platform Standard, App Test Platform Enterprise.
If the required entitlements are not enabled for your organization, the corresponding guardrail options will not appear in the UI. If you have already configured guardrails and the necessary entitlements are later disabled, your agents will simply skip those guardrails during execution, they will not cause agent runs to fail.
Access to out-of-the-box guardrails also depends on the cloud platform you use. For details, refer to Agents feature availability.
You can set up and configure out-of-the-box guardrails directly from your agent’s settings. The configuration applies automatically at runtime, based on the selected scope and action type.
- Open the Agent settings.
- In Studio Web, open your agent.
- Open the Agent settings panel.
- Go to the Guardrails tab and select Add guardrails.
- Choose a guardrail type from the available predefined guardrails:
- PII detection – Identifies and blocks sensitive information such as email addresses or physical addresses. This guardrail uses Azure Cognitive Services.
- Prompt injection – Detects and blocks malicious or manipulative prompts during LLM interactions. This guardrail uses Noma Security. Noma services are hosted in the United States, so data processed by these two guardrails may be handled outside your tenant’s region.
- Harmful content – Detects hate speech, self-harm, sexual content, and violence in LLM interactions. This guardrail uses Azure AI Content Safety.
- Intellectual property protection – Detects intellectual property leakage in model-generated text and code. This guardrail uses Azure AI Content Safety.
- User prompt attacks – Detects and blocks user prompt attacks, such as jailbreak or prompt injections. This guardrail uses Azure AI Content Safety.
- LLM as Judge – Evaluates prompts, responses, or LLM calls against custom instructions that you write in plain natural language, using the LLM you select as the judge model.
Configuring a PII detection guardrail
PII detection identifies and flags sensitive personal information — such as email addresses, phone numbers, and physical addresses — across agent interactions. You can apply it to agent prompts, LLM calls, or tool data. This guardrail uses Azure Cognitive Services.
- Add the PII detection guardrail to your agent.
- Fill in the following fields:
- Guardrail name – Enter a descriptive name for this guardrail.
- Guardrail description – Explain what it detects or where it applies.
- Select entities to detect. From the Entities to detect dropdown, choose the types of information you want to monitor (for example, email, phone, or address).
- Set a detection threshold. For each selected entity, define a Detection threshold between 0 and 1. A higher threshold makes detection stricter (fewer false positives), while a lower threshold makes it more sensitive.
- Choose the scopes. Select where you want the guardrail to apply:
- Agent – Checks the agent’s user prompt and output.
- LLM calls – Checks LLM requests and responses.
- Tools – Checks tool input and output data. You can select one or more scopes to apply the same detection logic across multiple stages of execution.
- If you select the Tools scope, choose one or more tools from the Select tools list. This lets you reuse the same guardrail across multiple tools in the same agent.
- Define the Action type. Configure how the system should respond when PII is detected:
- Log – Records the event without interrupting agent execution. The Severity level sets the importance level for the log entry:
- Info – For general information or low-impact findings.
- Warning – For potential risks that don’t block execution.
- Error – For critical detections that require review.
- Block – Stops the agent or tool execution when the guardrail is triggered.
- Block reason – Provide a brief explanation for blocking the action (for example, “Detected PII data in tool output”).
- Escalate – Sends an escalation when a violation occurs.
- Assign app to – Choose the escalation target type: a specific user, a defined user group, or an external address.
- Recipient – Search and select the recipient (name or email).
- Action app – Select the application that will handle the escalation.
- Log – Records the event without interrupting agent execution. The Severity level sets the importance level for the log entry:
- Enable for evaluations. Toggle Enable guardrail for evaluations to run this guardrail during agent testing or evaluation.
- Save the guardrail. Once configured, the guardrail automatically monitors all LLM requests and responses and blocks execution when prompt injection attempts are detected.
Configuring a Prompt injection guardrail
Prompt injection detection monitors LLM interactions for malicious or manipulative instructions that attempt to override the agent's intended behavior. This guardrail uses Noma Security.
- Add the Prompt injection guardrail to your agent.
- Fill in the following fields:
- Guardrail name – Enter a descriptive name for this guardrail.
- Guardrail description – Optionally explain what it detects or where it applies.
- Set the detection threshold. Specify the sensitivity level (for example, 0.8). Higher values make detection stricter and reduce false positives.
- Define the Action type. Configure how the system should handle detection events:
- Log – Records the event without interrupting agent execution. The Severity level sets the importance level for the log entry:
- Info – For general information or low-impact findings.
- Warning – For potential risks that don’t block execution.
- Error – For critical detections that require review.
- Block – Stops the agent or tool execution when the guardrail is triggered.
- Block reason – Provide a brief explanation for blocking the action.
- Escalate – Sends an escalation when a violation occurs.
- Assign app to – Choose the escalation target type: a specific user, a defined user group, or an external address.
- Recipient – Search and select the recipient (name or email).
- Action app – Select the application that will handle the escalation.
- Log – Records the event without interrupting agent execution. The Severity level sets the importance level for the log entry:
- Enable for evaluations. Toggle Enable guardrail for evaluations to run this guardrail during agent testing or evaluation.
- Save the guardrail. Once configured, the guardrail automatically monitors all LLM requests and responses and blocks execution when prompt injection attempts are detected.
Configuring a Harmful content guardrail
Harmful content detection identifies hate speech, self-harm, sexual content, and violence in LLM interactions. You can apply it before the LLM call, after, or at both stages. This guardrail uses Azure AI Content Safety.
- Add the Harmful content guardrail to your agent.
- Define the guardrail details. Fill in the following fields:
- Guardrail name – Enter a descriptive name for this guardrail.
- Guardrail description – Explain what it detects or where it applies.
- Select categories to detect. From the Entities to detect list, choose one or more types of content to monitor: hate speech, self-harm, sexual content, or Violence.
- Set a detection threshold. For each selected entity, define a Detection threshold between 0 and 6. Higher values require more severe content before triggering.
- Choose the scopes. Select where you want the guardrail to apply:
- Agent – Checks the agent’s user prompt and output.
- LLM calls – Checks LLM requests and responses.
- Tools – Checks tool input and output. You can select one or more scopes to apply the same detection logic across multiple stages of execution.
- If you select the Tools scope, choose one or more tools from the Select tools list. This lets you reuse the same guardrail across multiple tools in the same agent.
- Define the Action type. Configure how the system should respond when harmful content is detected:
- Log – Records the event without interrupting agent execution. The Severity level sets the importance level for the log entry:
- Info – For general information or low-impact findings.
- Warning – For potential risks that don't block execution.
- Error – For critical detections that require review.
- Block – Stops the agent or tool execution when the guardrail is triggered.
- Block reason – Provide a brief explanation for blocking the action.
- Escalate – Sends an escalation when a violation occurs.
- Assign app to – Choose the escalation target type: a specific user, a defined user group, or an external address.
- Recipient – Search and select the recipient (name or email).
- Action app – Select the application that will handle the escalation.
- Log – Records the event without interrupting agent execution. The Severity level sets the importance level for the log entry:
- Enable for evaluations. Toggle Enable guardrail for evaluations to run this guardrail during agent testing or evaluation.
- Save the guardrail. Once configured, the guardrail monitors LLM interactions and responds according to the configured action when harmful content is detected.
Configuring an Intellectual Property Protection guardrail
IP protection detects intellectual property leakage in model-generated text and code. This guardrail applies post-execution only, checking the model's response after the LLM call. This guardrail uses Azure AI Content Safety.
- Add the IP protection guardrail to your agent.
- Define the guardrail details. Fill in the following fields:
- Guardrail name – Enter a descriptive name for this guardrail.
- Guardrail description – Optionally explain what it detects or where it applies.
- From the Entities to detect list, choose what type of content to monitor: Text, Code, or both.
- Choose the scopes. Select where you want the guardrail to apply:
- Agent – Checks the agent’s input and output prompts.
- LLM calls – Checks the requests and responses exchanged with the model.
- Define the Action type. Configure how the system should handle detection events:
- Log – Records the event without interrupting agent execution. The Severity level sets the importance level for the log entry:
- Info – For general information or low-impact findings.
- Warning – For potential risks that don't block execution.
- Error – For critical detections that require review.
- Block – Stops the agent or tool execution when the guardrail is triggered.
- Block reason – Provide a brief explanation for blocking the action.
- Escalate – Sends an escalation when a violation occurs.
- Assign app to – Choose the escalation target type: a specific user, a defined user group, or an external address.
- Recipient – Search and select the recipient (name or email).
- Action app – Select the application that will handle the escalation.
- Log – Records the event without interrupting agent execution. The Severity level sets the importance level for the log entry:
- Enable for evaluations. Toggle Enable guardrail for evaluations to run this guardrail during agent testing or evaluation.
- Save the guardrail. Once configured, the guardrail monitors all LLM responses and responds according to the configured action when IP leakage is detected.
Configuring a User prompt attacks guardrail
User prompt attacks detects prompt injection and jailbreak attempts before they reach the LLM. This guardrails applies pre-execution only.
- Add the User prompt attacks guardrail to your agent.
- Define the guardrail details. Fill in the following fields:
- Guardrail name – Enter a descriptive name for this guardrail.
- Guardrail description – Explain what it detects or where it applies.
- Define the Action type. Configure how the system should respond when attacks are detected:
- Log – Records the event without interrupting agent execution. The Severity level sets the importance level for the log entry:
- Info – For general information or low-impact findings.
- Warning – For potential risks that don't block execution.
- Error – For critical detections that require review.
- Block – Stops the agent or tool execution when the guardrail is triggered.
- Block reason – Provide a brief explanation for blocking the action.
- Escalate – Sends an escalation when a violation occurs.
- Assign app to – Choose the escalation target type: a specific user, a defined user group, or an external address.
- Recipient – Search and select the recipient (name or email).
- Action app – Select the application that will handle the escalation.
- Log – Records the event without interrupting agent execution. The Severity level sets the importance level for the log entry:
- Enable for evaluations. Toggle Enable guardrail for evaluations to run this guardrail during agent testing or evaluation.
- Save the guardrail. Once configured, the guardrail automatically monitors all LLM requests and responses and blocks execution when user prompt attacks are detected.
Configuring an LLM as Judge guardrail
LLM as Judge evaluates agent prompts, responses, or LLM calls against custom instructions that you write in plain natural language, using an LLM as the judge. Unlike the other out-of-the-box guardrails, which check for a fixed, predefined category through a managed classification service, LLM as Judge lets you enforce free-form policies that don't map to an existing category, such as "never quote pricing" or "tool inputs must not contain internal code names."
Each check this guardrail performs makes a real LLM call, adding the judge model's token usage on top of the agent's own usage. This happens every time the guardrail runs, including during agent testing and evaluation runs when enabled for evaluations. You can choose a cheaper judge model where possible, because a judge doesn't need the agent model's capability to stay accurate.
- Add the LLM as Judge guardrail to your agent.
- Define the guardrail details. Fill in the following fields:
- Guardrail name – Enter a descriptive name for this guardrail.
- Guardrail description – Explain what it detects or where it applies.
- In the Rule prompt field, describe the policy to enforce in plain text, up to 4,000 characters, for example, "Never quote pricing" or "Always include this disclaimer."
- From the dedicated judge model picker, select the LLM to use as the Judge model, separate from the agent's own model list. The picker includes UiPath-managed models, and your own LLM Configurations if your organization uses bring your own model/subscription in AI Trust Layer.
- Optionally, provide up to two positive and two negative examples to improve judge accuracy. Each example is limited to 1,000 characters.
- Set a Detection threshold. This threshold defines the sensitivity level the judge uses to decide whether content violates the rule prompt.
- Choose where the guardrail applies:
- Agent – Checks the agent's user prompt and output.
- LLM calls – Checks LLM requests and responses.
- Tools – Checks tool input and output data. The guardrail applies before and after execution for the selected scope.
- Choose the Action type — how the guardrail responds when the judge returns a failing verdict: Log, Block, or Skip.
- Toggle Enable guardrail for evaluations to run this guardrail during agent testing and evaluation.
- Save the guardrail.
At runtime, the judge model returns a pass or fail verdict together with a human-readable reason, both visible in the trace details. Because the judge's own request and response are themselves an LLM call, centralized guardrails enabled in your AI Trust Layer policy for harmful content, prompt injection, PII, and IP protection also apply to them, the same as they apply to other LLM calls in the platform. For details, refer to Monitoring guardrails and Centralized guardrails.