FAIRBench Architecture¶
Overview¶
FAIRBench supports two evaluation modalities — text generation and image generation — that share a common scenario format, counterfactual expansion logic, and six fairness metrics. The modalities differ in how outputs are generated and analysed; everything downstream (metrics, scorecards, HTML reports) is identical.
Text Pipeline¶
Components¶
| Component | File | Role |
|---|---|---|
FairBenchEngine |
core/engine.py |
Orchestrates the full evaluation loop |
ModelAdapter |
adapters/base.py |
Abstract interface for LLM APIs |
AnthropicAdapter |
adapters/anthropic.py |
Claude (claude-sonnet-4-6, etc.) |
OpenAIAdapter |
adapters/openai.py |
GPT-4o, GPT-4, etc. |
OpenAICompatibleAdapter |
adapters/openai_compatible.py |
Together, Groq, Ollama, Mistral |
EvaluationPipeline |
evaluation/pipeline.py |
Generates + evaluates outputs |
DemographicClassifier |
evaluation/demographic.py |
Pronoun/name → demographic signal |
ToxicityEvaluator |
evaluation/toxicity.py |
Detoxify-based toxicity scores |
SentimentEvaluator |
evaluation/sentiment.py |
Positive/negative/neutral scores |
EmbeddingEvaluator |
evaluation/embeddings.py |
Sentence-transformer embeddings |
RefusalClassifier |
evaluation/refusal.py |
Rule-based refusal detection |
LLMJudgeEvaluator |
evaluation/llm_judge.py |
Optional Layer 2 LLM judge |
TriageRouter |
evaluation/triage.py |
Layer 3: flag outputs for review |
Execution flow¶
FairBenchEngine.evaluate()is called with a model, scenario set, and metric listScenarioRegistryloads scenario YAML filesCounterfactualGeneratorexpands each scenario into base + counterfactual promptsEvaluationPipelinegenerates outputs concurrently (rate-limited viaasyncio.Semaphore)- Layer 1 evaluators run on each output: embeddings, toxicity, sentiment, demographic signals, refusals
- Optional Layer 2 LLM judge evaluates semantic fairness
- Six fairness metrics are computed from
EvaluatedOutputobjects - Layer 3
TriageRouterflags severe cases for human review - Results are persisted to SQLite and returned as
EvaluationRun
Three-layer evaluator stack¶
Image Pipeline¶
Components¶
| Component | File | Role |
|---|---|---|
ImageBenchEngine |
core/image_engine.py |
Orchestrates image evaluation |
ImageModelAdapter |
adapters/image/base.py |
Abstract interface for image models |
DALLEAdapter |
adapters/image/dalle.py |
OpenAI gpt-image-1 |
StableDiffusionAdapter |
adapters/image/stable_diffusion.py |
HuggingFace Inference API or local diffusers |
VisionAnalyzer |
evaluation/image/vision_analyzer.py |
Claude Vision → structured ImageAnalysis |
CLIPEvaluator |
evaluation/image/clip_evaluator.py |
CLIP embeddings + text-image similarity probes |
EvaluatedImage |
core/image_types.py |
Image with analysis annotations |
Execution flow¶
- Scenarios are loaded in the same YAML format as text scenarios
CounterfactualGeneratorexpands prompts (identical to text pipeline)ImageModelAdapter.generate()is called concurrently for each promptVisionAnalyzercalls Claude Vision on each image →ImageAnalysis(gender, skin tone, setting, equipment, stereotypes)CLIPEvaluatorcomputes visual embeddings and text-image similarity probesEvaluatedImage.to_evaluated_output()bridges image signals intoEvaluatedOutput- The six fairness metrics run unchanged on bridged
EvaluatedOutputobjects - Scorecard is generated with image-specific fields (gender distribution, setting distribution, stereotype list)
Bridge: image → text metric space¶
The bridge maps image analysis signals into the same fields the text metrics consume:
| Image signal | EvaluatedOutput field |
Used by metric |
|---|---|---|
| CLIP visual embedding | embedding |
CDS (cosine distance), ODE (diversity) |
counterfactual_value (from scenario) |
counterfactual_value |
RSI (distribution counting) |
stereotype_severity |
toxicity.identity_attack |
HSI (harm scoring) |
image_quality_score / 10 |
helpfulness_score |
DSI (quality disparity) |
is_refused |
is_refusal |
DSI (refusal rate) |
perceived_gender, skin_tone_label |
detected_entities |
RSI "detected" mode |
This bridge means the six metrics and the scorecard/HTML report work identically for both modalities with no duplication.
Scenario format¶
Both modalities use the same YAML format:
name: my_scenarios
version: "1.0"
description: "…"
dimensions:
- representational
- distributional
scenarios:
- id: example_scenario
prompt: "A software engineer solves a difficult problem."
counterfactuals:
- attribute: gender
variants:
- prompt: "A female software engineer solves a difficult problem."
value: female
- prompt: "A male software engineer solves a difficult problem."
value: male
- attribute: race
variants:
- prompt: "A Black software engineer solves a difficult problem."
value: black
Built-in text scenarios live in src/fairbench_genai/scenarios/builtin/.
Built-in image scenarios live in src/fairbench_genai/scenarios/image/.
Storage¶
Evaluation runs are persisted to SQLite (~/.fairbench/fairbench.db by default) via SQLiteBackend. The schema stores runs, per-output annotations, and metric results. Runs can be retrieved by ID with fairbench show <run_id> or via engine.get_run(run_id).
Image pipeline runs are not yet persisted to SQLite (results are returned in-memory as ImageEvaluationRun).
Extension points¶
Custom model adapter (text)¶
from fairbench_genai.adapters.base import ModelAdapter
from fairbench_genai.core.types import GeneratedOutput, ModelInfo
class MyAdapter(ModelAdapter):
async def generate(self, prompt, config=None) -> GeneratedOutput: ...
def get_model_info(self) -> ModelInfo: ...
@property
def name(self) -> str: return "my-model"
engine.register_adapter("my-model", MyAdapter())
Custom image adapter¶
from fairbench_genai.adapters.image.base import ImageModelAdapter
from fairbench_genai.core.image_types import GeneratedImage, ModelInfo
class MyImageAdapter(ImageModelAdapter):
async def generate(self, prompt, config=None) -> GeneratedImage: ...
def get_model_info(self) -> ModelInfo: ...
@property
def name(self) -> str: return "my-image-model"
Custom metric¶
from fairbench_genai.metrics.base import Metric
from fairbench_genai.core.types import EvaluatedOutput, MetricResult
class MyMetric(Metric):
def compute(self, outputs: list[EvaluatedOutput], baseline=None) -> MetricResult: ...
def interpret(self, result) -> str: ...
@property
def name(self) -> str: return "MY"
@property
def description(self) -> str: return "…"
engine.register_metric(MyMetric())
Custom scenarios¶
from fairbench_genai.core.types import Scenario, CounterfactualGroup, CounterfactualVariant
scenario = Scenario(
id="my_scenario",
prompt="Describe a surgeon.",
counterfactuals=[
CounterfactualGroup(attribute="gender", variants=[
CounterfactualVariant(prompt="Describe a female surgeon.", attribute_value="female"),
CounterfactualVariant(prompt="Describe a male surgeon.", attribute_value="male"),
])
],
)
result = await engine.evaluate(model=adapter, scenarios=[scenario])