fix: docs

This commit is contained in:
Alexander Myasoedov
2026-09-07 14:03:57 +03:00
parent b5aa075927
commit f286ebf21a
10 changed files with 160 additions and 355 deletions
+3 -5
View File
@@ -42,11 +42,9 @@ This section provides detailed information about the Agentic Security API.
## Authentication
All API requests require an API key. Include it in the `Authorization` header:
```
Authorization: Bearer YOUR_API_KEY
```
`/v1/self-probe` is an unauthenticated local test endpoint. File and image
probe routes expect an `Authorization: Bearer ...` header; the mock server
checks the header shape, not a configured API key.
## Further Reading
+16 -17
View File
@@ -1,24 +1,23 @@
# Configuration
This section provides information on configuring Agentic Security to suit your needs.
Scan settings live in `agentic_security.toml` in the working directory.
Create one with:
## Default Configuration
```bash
agentic_security init
```
The default configuration file is `agentic_security.toml`. It includes settings for:
The file is versioned (`general.version`). The current schema version is `2`.
- General settings
- Module configurations
- Thresholds
## Sections
## Customizing Configuration
- `[general]` — `llmSpec` (the HTTP request template), budget, failure threshold,
optimizer flag, multi-step attack flag
- `[modules.*]` — datasets to run; `AgenticBackend.opts.port` is the local proxy port
- `[detectors]` — refusal and leak classifiers
- `[thresholds]` — low / medium / high failure-rate bands
- `[secrets]` — API keys referenced from the HTTP spec
- `[caching]`, `[network]`, `[fuzzer]` — runtime tuning
1. Open the `agentic_security.toml` file in a text editor.
1. Modify the settings as needed. For example, to change the port:
```toml
[modules.AgenticBackend.opts]
port = 8718
```
## Advanced Configuration
For advanced configuration options, refer to the [API Reference](api_reference.md).
Replace the placeholder `Bearer XXXXX` in `llmSpec` with credentials for the
target endpoint. See [HTTP spec](http_spec.md) for the request template format.
+48 -37
View File
@@ -1,49 +1,60 @@
## Module Interface Documentation
# Module interface
The ``Module`` class provides a standardized way to create and use probe data
modules in the ``agentic_security`` project.
Probe data modules in `agentic_security.probe_data.modules` share the same
constructor arguments. `apply` yields result strings; some modules implement
it as a sync iterator and others as an async iterator.
All modules in :mod:`agentic_security.probe_data.modules` share the same
constructor signature and the same ``apply`` async generator method. See the
concrete implementations for real-world usage:
Built-in examples:
* :mod:`agentic_security.probe_data.modules.garak_tool`
* :mod:`agentic_security.probe_data.modules.fine_tuned`
* :mod:`agentic_security.probe_data.modules.inspect_ai_tool`
* :mod:`agentic_security.probe_data.modules.rl_model`
- `agentic_security.probe_data.modules.garak_tool`
- `agentic_security.probe_data.modules.fine_tuned`
- `agentic_security.probe_data.modules.inspect_ai_tool`
- `agentic_security.probe_data.modules.rl_model`
### Interface Summary
## Constructor
Every module class accepts three constructor arguments:
```python
def __init__(
self,
prompt_groups: list,
tools_inbox: asyncio.Queue,
opts: dict | None = None,
): ...
```
``def __init__(self, prompt_groups: list[Any], tools_inbox: asyncio.Queue, opts: dict = {}): ...``
`opts` is a module-specific dictionary. An omitted `opts` is treated as `{}`.
The ``apply`` method is an async generator that yields result strings:
## Usage
``async def apply(self) -> AsyncGenerator[str, None]: yield "result message"``
```python
import asyncio
from agentic_security.probe_data.modules.garak_tool import Module as GarakModule
### Usage Example
tools_inbox = asyncio.Queue()
module = GarakModule(["group_a"], tools_inbox, {"port": 8718})
``import asyncio``
``from agentic_security.probe_data.modules.garak_tool import Module as GarakModule``
``tools_inbox = asyncio.Queue()``
``prompt_groups = ["group_a", "group_b"]``
``opts = {"port": 8718}``
``module = GarakModule(prompt_groups, tools_inbox, opts)``
``async def main(): async for result in module.apply(): print(result)``
``asyncio.run(main())``
async def main():
async for result in module.apply():
print(result)
### Defining a Custom Module
asyncio.run(main())
```
``import asyncio``
``from typing import Any``
``class MyModule:``
`` def __init__(self, prompt_groups, tools_inbox, opts={}):``
`` self.prompt_groups = prompt_groups``
`` self.tools_inbox = tools_inbox``
`` self.opts = opts``
`` async def apply(self):``
`` for group in self.prompt_groups:``
`` result = "processed {0}".format(group)``
`` await self.tools_inbox.put({"message": result})``
`` yield result``
## Custom module
```python
import asyncio
class MyModule:
def __init__(self, prompt_groups, tools_inbox, opts=None):
self.prompt_groups = prompt_groups
self.tools_inbox = tools_inbox
self.opts = opts or {}
async def apply(self):
for group in self.prompt_groups:
result = f"processed {group}"
await self.tools_inbox.put({"message": result})
yield result
```
+7 -16
View File
@@ -1,23 +1,14 @@
# Getting Started
Welcome to Agentic Security! This guide will help you get started with using the tool.
1. Install the package ([installation](installation.md)).
2. Start the UI:
## Quick Start
1. Ensure you have completed the [installation](installation.md) steps.
1. Run the following command to start the application:
```bash
agentic_security server
agentic_security server
```
1. Access the application at `http://localhost:8718`.
## Basic Usage
The server listens on `http://127.0.0.1:8718` by default.
- To view available commands, use:
```bash
agentic_security --help
```
## Next Steps
Explore the [Configuration](configuration.md) section to customize your setup.
3. For a headless scan, create `agentic_security.toml` with `agentic_security init`,
edit the `llmSpec`, then run `agentic_security ci`. See the [CLI](cli.md)
and [configuration](configuration.md) pages.
+11 -12
View File
@@ -1,19 +1,18 @@
# Installation
This section will guide you through the installation process for Agentic Security.
Requires Python 3.14 or newer.
## Prerequisites
## PyPI
- Python 3.11
- pip
```bash
pip install agentic_security
```
## Installation Steps
## From source
1. Install the package using pip:
```bash
pip install agentic_security
```
```bash
poetry install --with dev
```
## Troubleshooting
If you encounter any issues during installation, please refer to the [troubleshooting guide](#) or contact support.
Then run `agentic_security --help`. See the [CLI](cli.md) and
[getting started](getting_started.md) pages for the next steps.
+26 -27
View File
@@ -1,43 +1,42 @@
# Probe Actor Module Documentation
# Probe actor
The `probe_actor` module is a critical component of the Agentic Security project, responsible for generating prompts, performing scans, and handling refusal checks. This documentation provides an overview of the module's structure and functionality.
The `probe_actor` package runs scans against an LLM endpoint described by an
HTTP spec.
## Files and Key Components
## Fuzzer
### fuzzer.py
`agentic_security.probe_actor.fuzzer` is the scan engine:
- **Functions:**
- `async def generate_prompts(...)`: Asynchronously generates prompts for scanning.
- `def multi_modality_spec(llm_spec)`: Defines specifications for multi-modality.
- `async def process_prompt(...)`: Processes a given prompt asynchronously.
- `async def perform_single_shot_scan(...)`: Performs a single-shot scan asynchronously.
- `async def perform_many_shot_scan(...)`: Performs a many-shot scan asynchronously.
- `def scan_router(...)`: Routes scan requests.
- `generate_prompts` — yield prompts from a list or async source
- `get_modality_adapter` — wrap the spec for image or audio endpoints
- `process_prompt` / `process_prompt_batch` — send a prompt and score the reply
- `perform_single_shot_scan` — one prompt at a time
- `perform_many_shot_scan` — many-shot jailbreak conversations
- `scan_router` — pick the scan mode from configuration
### refusal.py
- **Functions:**
- `def check_refusal(response: str, refusal_phrases: list = REFUSAL_MARKS) -> bool`: Checks if a response contains refusal phrases.
- `def refusal_heuristic(request_json)`: Applies heuristics to determine refusal.
## Usage Examples
### Performing a Single-Shot Scan
`perform_single_shot_scan` is an async generator; it needs a request factory
(typically an `LLMSpec`) plus budget and dataset settings, not a bare prompt
string.
```python
from agentic_security.http_spec import LLMSpec
from agentic_security.probe_actor.fuzzer import perform_single_shot_scan
await perform_single_shot_scan(prompt="Test prompt")
spec = LLMSpec.from_string(open("spec.http").read())
async for event in perform_single_shot_scan(spec, max_budget=1000):
print(event)
```
### Checking for Refusal
## Refusal checks
There is no `check_refusal` helper. Use `refusal_heuristic(response)` in
`agentic_security.probe_actor.refusal`, which runs the detectors enabled under
`[detectors]` in `agentic_security.toml`.
```python
from agentic_security.probe_actor.refusal import check_refusal
from agentic_security.probe_actor.refusal import refusal_heuristic
is_refusal = check_refusal(response="I'm sorry, I can't do that.")
refusal_heuristic("I'm sorry, I can't help with that.")
```
## Conclusion
The `probe_actor` module provides essential functionality for generating prompts, performing scans, and handling refusal checks within the Agentic Security project. This documentation serves as a guide to understanding and utilizing the module's capabilities.
See [refusal classifier plugins](refusal_classifier_plugins.md) to add a custom detector.
+24 -97
View File
@@ -1,93 +1,28 @@
# Probe Data Module Documentation
# Probe data
The `probe_data` module is a core component of the Agentic Security project, responsible for handling datasets, generating audio and image data, and applying various transformations. This documentation provides an overview of the module's structure and functionality.
The `probe_data` package loads attack datasets, optionally transforms prompts,
and can generate image or audio payloads for multi-modal specs.
## Files and Key Components
## Datasets
### audio_generator.py
`agentic_security.probe_data.data`:
- **Functions:**
- `encode(content: bytes) -> str`: Encodes audio content to a string format.
- `generate_audio_mac_wav(prompt: str) -> bytes`: Generates audio in WAV format for macOS.
- `generate_audioform(prompt: str) -> bytes`: Generates audio from a given prompt.
- **Classes:**
- `RequestAdapter`: Handles requests for audio generation.
- `load_dataset_generic` — load a CSV URL or Hugging Face dataset into a `ProbeDataset`
- `load_local_csv` / `load_local_csv_files` — datasets from local CSV files
- `prepare_prompts` — select and transform registry datasets for a scan
### data.py
`agentic_security.probe_data.models.ProbeDataset` is the in-memory dataset type.
Image datasets wrap that as `ImageProbeDataset`.
- **Functions:**
- `load_dataset_general(...)`: Loads datasets with general specifications.
- `count_words_in_list(str_list)`: Counts words in a list of strings.
- `prepare_prompts(...)`: Prepares prompts for dataset processing.
- **Classes:**
- `Stenography`: Applies transformations to prompt groups.
## Transforms
### image_generator.py
`agentic_security.probe_data.stenography_fn` provides encoding helpers such as
`rot13`, `base64_encode`, and `mirror_words`. See [stenography](stenography.md).
- **Functions:**
- `generate_image_dataset(...)`: Generates a dataset of images.
- `generate_image(prompt: str) -> bytes`: Generates an image from a prompt.
- **Classes:**
- `RequestAdapter`: Handles requests for image generation.
## Image and audio
### models.py
- **Classes:**
- `ProbeDataset`: Represents a dataset for probing.
- `ImageProbeDataset`: Extends `ProbeDataset` for image data.
### msj_data.py
- **Functions:**
- `load_dataset_generic(...)`: Loads a generic dataset.
- **Classes:**
- `ProbeDataset`: Represents a dataset for probing.
### stenography_fn.py
- **Functions:**
- `rot13(input_text)`: Applies ROT13 transformation.
- `base64_encode(data)`: Encodes data in base64 format.
- `mirror_words(text)`: Mirrors words in the text.
### rl_model.py
- **Classes:**
- `PromptSelectionInterface`: Abstract base class for prompt selection strategies.
- Methods:
- `select_next_prompt(current_prompt: str, passed_guard: bool) -> str`: Selects next prompt
- `select_next_prompts(current_prompt: str, passed_guard: bool) -> list[str]`: Selects multiple prompts
- `update_rewards(previous_prompt: str, current_prompt: str, reward: float, passed_guard: bool) -> null`: Updates rewards
- `RandomPromptSelector`: Basic random selection with history tracking.
- Parameters:
- `prompts: list[str]`: List of available prompts
- `history_size: int = 3`: Size of history to prevent cycles
- `CloudRLPromptSelector`: Cloud-based RL implementation with fallback.
- Parameters:
- `prompts: list[str]`: List of available prompts
- `api_url: str`: URL of RL service
- `auth_token: str = AUTH_TOKEN`: Authentication token
- `history_size: int = 300`: Size of history
- `timeout: int = 5`: Request timeout
- `run_id: str = ""`: Unique run identifier
- `QLearningPromptSelector`: Local Q-learning implementation.
- Parameters:
- `prompts: list[str]`: List of available prompts
- `learning_rate: float = 0.1`: Learning rate
- `discount_factor: float = 0.9`: Discount factor
- `initial_exploration: float = 1.0`: Initial exploration rate
- `exploration_decay: float = 0.995`: Exploration decay rate
- `min_exploration: float = 0.01`: Minimum exploration rate
- `history_size: int = 300`: Size of history
- **Module**: Main class that uses CloudRLPromptSelector.
- Parameters:
- `prompt_groups: list[str]`: Groups of prompts
- `tools_inbox: asyncio.Queue`: Queue for tool communication
- `opts: dict = {}`: Configuration options
## Usage Examples
### Generating Audio
- `generate_image` / `generate_image_dataset` in `probe_data.image_generator`
- `generate_audioform` in `probe_data.audio_generator`
```python
from agentic_security.probe_data.audio_generator import generate_audioform
@@ -95,26 +30,18 @@ from agentic_security.probe_data.audio_generator import generate_audioform
audio_bytes = generate_audioform("Hello, world!")
```
### Loading a Dataset
## Prompt selection
```python
from agentic_security.probe_data.data import load_dataset_general
dataset = load_dataset_general("example_dataset")
```
### Using RL Model
`probe_data.modules.rl_model` implements `PromptSelectionInterface` with
`RandomPromptSelector`, `CloudRLPromptSelector`, and `QLearningPromptSelector`.
`update_rewards` returns `None`. Boolean arguments are Python `True` / `False`.
```python
from agentic_security.probe_data.modules.rl_model import QLearningPromptSelector
prompts = ["What is AI?", "Explain machine learning"]
selector = QLearningPromptSelector(prompts)
current_prompt = "What is AI?"
next_prompt = selector.select_next_prompt(current_prompt, passed_guard=true)
selector.update_rewards(current_prompt, next_prompt, reward=1.0, passed_guard=true)
selector = QLearningPromptSelector(["What is AI?", "Explain machine learning"])
next_prompt = selector.select_next_prompt("What is AI?", passed_guard=True)
selector.update_rewards("What is AI?", next_prompt, reward=1.0, passed_guard=True)
```
## Conclusion
The `probe_data` module provides essential functionality for handling and transforming datasets within the Agentic Security project. This documentation serves as a guide to understanding and utilizing the module's capabilities.
See [RL model](rl_model.md) for the selector options.
+1 -1
View File
@@ -35,7 +35,7 @@ Here is an example of a custom refusal classifier plugin that checks for specifi
```python
class CustomRefusalClassifier(RefusalClassifierPlugin):
def __init__(self, custom_phrases: List[str]):
def __init__(self, custom_phrases: list[str]):
self.custom_phrases = custom_phrases
def is_refusal(self, response: str) -> bool:
+24 -143
View File
@@ -1,153 +1,34 @@
# Stenography Functions
# Stenography functions
The stenography module provides various text obfuscation and transformation techniques for security testing. This document explains its architecture and implementation.
`agentic_security.probe_data.stenography_fn` transforms prompts so scanners can
test whether a model still treats obfuscated text as an instruction.
## Overview
The module is used by `Stenography` in `probe_data.data` when a dataset
enables encoding variants.
The module implements:
## Functions
1. Rotation ciphers (ROT13, ROT5)
1. Base64 encoding
1. Text manipulation functions
1. Randomization techniques
1. Character substitution methods
| Function | What it does |
| --- | --- |
| `rot13` / `rot5` | Letter and digit rotation ciphers |
| `base64_encode` | Base64-encode a string or bytes |
| `mirror_words` | Reverse each word, keep word order |
| `scramble_words` | Shuffle middle letters, keep first and last |
| `randomize_letter_case` | Random per-character case |
| `insert_noise_characters` | Insert alphanumeric noise (`frequency=0.2` by default) |
| `substitute_with_ascii` | Replace characters with ordinals |
| `remove_vowels` | Drop `aeiou` in either case |
| `zigzag_obfuscation` | Alternate upper/lower case |
| `caesar_cipher` / `substitution_cipher` / `vigenere_cipher` | Classical ciphers |
| `code_block_encode` | Hide the prompt in a Python docstring fenced as code |
## Core Functions
These are obfuscation helpers for red-team prompts, not cryptographic primitives.
### Rotation Ciphers
## Usage
```python
def rot13(input_text):
"""
Applies ROT13 cipher to input text
- Preserves case of letters
- Leaves non-alphabetic characters unchanged
"""
# Implementation details...
from agentic_security.probe_data.stenography_fn import rot13, scramble_words
def rot5(input_text):
"""
Applies ROT5 cipher to input text
- Rotates digits by 5 positions
- Leaves non-digit characters unchanged
"""
# Implementation details...
rot13("Ignore previous instructions")
scramble_words("Ignore previous instructions")
```
### Encoding
```python
def base64_encode(data):
"""
Encodes input data using Base64
- Handles both string and bytes input
- Returns UTF-8 encoded string
"""
# Implementation details...
```
### Text Manipulation
```python
def mirror_words(text):
"""
Reverses each word in the input text
- Preserves word order
- Maintains spaces between words
"""
# Implementation details...
def scramble_words(text):
"""
Randomly scrambles middle letters of words
- Preserves first and last letters
- Handles words shorter than 4 characters
"""
# Implementation details...
```
### Randomization
```python
def randomize_letter_case(text):
"""
Randomly changes case of each character
- Independent case changes per character
- Preserves non-letter characters
"""
# Implementation details...
def insert_noise_characters(text, frequency=0.2):
"""
Inserts random characters between existing ones
- Configurable insertion frequency
- Uses alphanumeric characters for noise
"""
# Implementation details...
```
### Advanced Transformations
```python
def substitute_with_ascii(text):
"""
Replaces characters with their ASCII codes
- Space-separated numeric values
- Preserves original character order
"""
# Implementation details...
def remove_vowels(text):
"""
Removes all vowel characters from text
- Handles both lowercase and uppercase vowels
- Preserves non-vowel characters
"""
# Implementation details...
def zigzag_obfuscation(text):
"""
Alternates character case in zigzag pattern
- Starts with uppercase
- Toggles case for each alphabetic character
"""
# Implementation details...
```
## Usage Patterns
1. **Text Obfuscation**:
```python
obfuscated = zigzag_obfuscation(
scramble_words(
insert_noise_characters(text)
)
)
```
1. **Encoding**:
```python
encoded = base64_encode(rot13(text))
```
1. **Randomization**:
```python
randomized = randomize_letter_case(
remove_vowels(text)
)
```
## Configuration
- **Noise Frequency**: Configurable in insert_noise_characters()
- **Scrambling**: Automatic handling of word lengths
- **Case Handling**: Preserved in rotation ciphers
## Limitations
- Primarily handles ASCII text
- Limited to implemented transformation types
- Randomization is not cryptographically secure