Add guardrails module: prompt injection defences (document filter, tool policy, output validation) measured against an always-obeying stub model

Co-Authored-By: Claude Sonnet 5.5 <[email protected]>
Claude-Session: https://claude.ai/code/session_01JXVi2GMQ7bR5EmbUFdDj7N
This commit is contained in:
Claude
2026-10-09 06:50:23 +00:00
parent cc2a1b8bcb
commit b0bba995e6
35 changed files with 1349 additions and 0 deletions
+12
View File
@@ -0,0 +1,12 @@
# Tool policy: one call per rule
refund 25 USD on the customer's order -> "refunded 25.0 on A-1001"
refund 400 USD on the customer's order -> DENIED: amount outside 0 < x <= 50 USD (needs a human)
refund 10 USD on someone else's order -> DENIED: order is not this customer's
email to [email protected] -> DENIED: recipient is not the signed-in customer
email to the signed-in customer -> "sent to [email protected]"
a 4th call in one request (limit is 3) -> DENIED: more than 3 tool calls in one request
sendEmail on a read-only endpoint -> (tool not exposed)
side effects that happened: refunds=[Refund[orderId=A-1001, amountUsd=25.0]] emails=[[email protected]]
denials recorded (all policies): [refund: amount outside 0 < x <= 50 USD (needs a human), refund: order is not this customer's, sendEmail: recipient is not the signed-in customer, lookupOrder: more than 3 tool calls in one request]