Add guardrails module: prompt injection defences (document filter, tool policy, output validation) measured against an always-obeying stub model

Co-Authored-By: Claude Sonnet 5.5 <[email protected]>
Claude-Session: https://claude.ai/code/session_01JXVi2GMQ7bR5EmbUFdDj7N
This commit is contained in:
Claude
2026-10-09 06:50:23 +00:00
parent cc2a1b8bcb
commit b0bba995e6
35 changed files with 1349 additions and 0 deletions
@@ -0,0 +1,8 @@
# Validating a typed model answer
valid ACCEPTED SHIPPING p2
category not in the enum rejected: not parseable as Triage: InvalidFormatException
priority out of range rejected: constraint violation: priority must be less than or equal to 5
reply too long rejected: constraint violation: reply size must be between 0 and 280
not JSON at all rejected: not parseable as Triage: StreamReadException
well-formed, hostile link rejected: reply text: link to evil.example