Python 3.10+ License: MIT Tests
Left: base LLM agents fail — firms crash prices into bankruptcy (B2C), and a Sybil principal floods the market with deceptive listings (C2C). Right: aligned agents restore equilibrium — Stabilizing Firms hold a price floor; Skeptical Guardians detect and reject the Sybil cluster.
Simulation architecture: B2C market (left) with Poisson consumer arrivals, firm pricing agents, and discovery limits; C2C market (right) with buyer/seller agents, Sybil principal, and reputation signals.
As AI agents increasingly operate autonomously in digital marketplaces, their collective behavior introduces systemic risks that standard alignment — targeting helpfulness, harmlessness, and honesty — does not address. An agent that is individually rational can drive a market into collapse through locally optimal but globally destructive decisions. We define Economic Alignment as the property of contributing to stable, fair markets rather than chaotic or exploitative ones, and introduce Agent Bazaar to benchmark it.
We study two canonical failure modes. THE_CRASH: in B2C markets, LLM firms engage in a destructive undercutting race until prices fall below unit cost, triggering mass bankruptcy — an LLM-native analog of the 2010 Flash Crash. THE_LEMON_MARKET: in C2C markets, a Sybil principal operates K coordinated seller identities, rotating them when reputation degrades to perpetuate fraud at scale, an amplified version of Akerlof's market for lemons.
For each failure mode, Agent Bazaar tests intervention mechanisms: Stabilizing Firms enforce a price floor against the undercutting spiral; Skeptical Guardians detect and reject deceptive listings. We evaluate frontier and open-weight models (3B–405B) across both scenarios and introduce the Economic Alignment Score (EAS), a unified scalar aggregating stability, integrity, welfare, and profitability into a single cross-model metric.
The simulator builds on infrastructure from LLM Economist and extends it with agent-agent goods trading in different market structures, firm/buyer/seller/Sybil agents, and the EAS metric.
conda create -n agent-bazaar python=3.12 -y
conda activate agent-bazaar# From PyPI
pip install agent-bazaar
# Or development install from source
git clone https://github.com/sethkarten/AI-Bazaar.git
cd AI-Bazaar
pip install -e .Set the keys for whichever LLM providers you plan to use:
# Google Gemini (Google AI Studio)
export GOOGLE_API_KEY="your_google_key"
# or for Vertex AI: set up Application Default Credentials
# gcloud auth application-default login
# OpenAI
export OPENAI_API_KEY="your_openai_key"
# Anthropic
export ANTHROPIC_API_KEY="your_anthropic_key"
# OpenRouter (access many models through one API)
export OPENROUTER_API_KEY="your_openrouter_key"Ollama runs models locally on your GPU without any API quota.
Install:
- Windows/macOS: download from ollama.com
- Linux:
curl -fsSL https://ollama.com/install.sh | sh
Pull a model and start the server:
ollama pull llama3.1:8b
ollama serve # leave this terminal openTo allow parallel requests (recommended for simulations; ensure sufficient GPU memory to hold 4 instances of the model before setting):
# Linux/macOS
export OLLAMA_NUM_PARALLEL=4 && ollama serve
# Windows (PowerShell)
$env:OLLAMA_NUM_PARALLEL = "4"; ollama servevLLM serves Hugging Face models via an OpenAI-compatible API. On Windows, vLLM works best in WSL2.
pip install vllm
# Start a server (run in a separate terminal)
python -m vllm.entrypoints.openai.api_server \
--model google/gemma-3-4b-it \
--port 8009For gated models, set your Hugging Face token first:
export HF_TOKEN="hf_..."
# or: huggingface-cli loginUse --llm google/gemma-3-4b-it --service vllm --port 8009 in your simulation command to point at this server.
# THE_CRASH: 5 LLM firms, 50 consumers, 20 timesteps (Gemini)
python -m agent_bazaar.main \
--consumer-scenario THE_CRASH \
--firm-type LLM --num-firms 5 --num-consumers 50 \
--use-cost-pref-gen --no-diaries --prompt-algo cot \
--llm gemini-2.5-flash --max-timesteps 20 --name crash_test
# LEMON_MARKET: 12 sellers (3 Sybil), 12 LLM buyers
python -m agent_bazaar.main \
--consumer-scenario LEMON_MARKET \
--num-sellers 12 --num-buyers 12 \
--sybil-cluster-size 3 --reputation-initial 0.8 \
--no-diaries --prompt-algo cot \
--llm gemini-2.5-flash --max-timesteps 20 --name lemon_test
# Local model via Ollama
python -m agent_bazaar.main \
--consumer-scenario THE_CRASH \
--firm-type LLM --num-firms 3 --num-consumers 20 \
--llm llama3.1:8b --service ollama --port 11434 \
--max-timesteps 10 --name local_testAfter running a simulation, inspect results in the Streamlit dashboard:
streamlit run agent_bazaar/viz/dashboard.pyState files are stored at logs/<run_name>/states.json. The dashboard lists all available runs and lets you explore per-timestep market state.
AI-Bazaar/
├── agent_bazaar/ # Main package
│ ├── agents/ # Agent implementations
│ │ ├── firm.py # LLM firm / Stabilizing firm
│ │ ├── buyer.py # LLM buyer / Skeptical Guardian
│ │ ├── seller.py # LLM seller / Sybil principal
│ │ └── consumer.py # CES consumer
│ ├── models/ # LLM provider integrations
│ │ ├── openai_model.py
│ │ ├── gemini_model.py
│ │ ├── vllm_model.py
│ │ ├── openrouter_model.py
│ │ └── base.py
│ ├── env/ # BazaarEnv simulation environment
│ ├── market_core/ # Market clearing and mechanics
│ ├── utils/ # Shared utilities
│ ├── viz/ # Streamlit dashboard
│ │ └── dashboard.py
│ └── main.py # Entry point and CLI
├── scripts/ # Experiment runners
│ ├── exp1.py # THE_CRASH main sweep
│ ├── exp1_eas_sweep.py # THE_CRASH × open-weight model sweep
│ ├── exp2.py # LEMON_MARKET main sweep
│ ├── exp2_2.py # LEMON_MARKET (no seller IDs ablation)
│ ├── exp2_eas_sweep.py # LEMON_MARKET × buyer model sweep
│ ├── exp3.py # Adversarial shock experiments
│ ├── exp3_open_weights_sweep.py
│ ├── exp5.py # Discovery limit firms (DLF) ablation
│ ├── exp6.py # Consumer procedural personas
│ ├── analyze_lemon_prompts.py
│ ├── compile_listing_corpus.py
│ ├── consolidate_states.py
│ └── extract_sybil_prompts.py
├── corpus/ # Pre-compiled listing corpus for LEMON_MARKET
├── documentation/ # Run commands and model reference
├── examples/ # Usage examples
├── experiments/ # Experiment runner framework
├── tests/ # Test suite
└── README.md
Run from the project root:
python -m agent_bazaar.main [OPTIONS]| Argument | Default | Description |
|---|---|---|
--firm-type |
FIXED |
LLM or FIXED firm agents |
--num-firms |
5 |
Number of firms (THE_CRASH) |
--num-consumers |
50 |
Number of consumers |
--num-sellers |
— | Alias for --num-firms in LEMON_MARKET |
--num-buyers |
— | Alias for --num-consumers in LEMON_MARKET |
--num-stabilizing-firms |
0 |
Number of Stabilizing Firms (THE_CRASH). First N LLM firms get the stabilizing prompt and enforce price ≥ unit cost |
--seller-type |
FIXED |
LEMON_MARKET: LLM generates descriptions; FIXED uses templates |
--sybil-cluster-size |
0 |
LEMON_MARKET: number of Sybil identities (last K of --num-sellers). 0 = no Sybil |
--firm-personas |
— | Comma-separated persona:count pairs for non-stabilizing firms (e.g. competitive:3,volume_seeker:2). Valid: competitive, volume_seeker, reactive, cautious |
--seller-personas |
— | LEMON_MARKET: comma-separated persona:count for honest sellers. Valid: standard, detailed, terse, optimistic |
--disable-firm-personas |
off | Strip behavioral archetypes from all firm prompts |
--enable-consumer-personas |
off | THE_CRASH: assign behavioral personas to CES consumers round-robin (LOYAL, SMALL_BIZ, REP_SEEKER, VARIETY) |
--unit-cost |
2.0 |
Unit cost of production |
--max-supply-unit-cost |
1.0 |
Upper bound on randomly drawn per-firm unit costs |
--firm-initial-cash |
500.0 |
Starting cash balance for each firm |
--overhead-costs |
14.0 |
Fixed overhead cost per timestep per firm |
--firm-markup |
0.50 |
FIXED firm: markup over unit cost |
--firm-tax-rate |
0.05 |
Tax rate on firm cash each timestep |
--use-cost-pref-gen |
off | Generate heterogeneous supply costs and CES preferences via the heterogeneity module |
--use-gen-ces |
off | Generate CES parameters via LLM for consumers |
| Argument | Default | Description |
|---|---|---|
--consumer-scenario |
RACE_TO_BOTTOM |
RACE_TO_BOTTOM, EARLY_BIRD, PRICE_DISCRIMINATION, RATIONAL_BAZAAR, BOUNDED_BAZAAR, THE_CRASH, LEMON_MARKET |
--consumer-type |
CES |
CES or FIXED consumer agents |
--max-timesteps |
100 |
Simulation length |
--discovery-limit-consumers |
3 |
Max firms a consumer polls before ordering (0 = no limit) |
--discovery-limit-firms |
0 |
Max competitors each firm observes (0 = no limit) |
--wtp-algo |
none |
none (always order), wtp (CES willingness-to-pay), ewtp (expected WTP) |
--poisson-demand-lambda |
— | Poisson arrival rate for consumer participation per step. Default: all consumers participate (THE_CRASH defaults to 0.6 × num_consumers) |
--info-asymmetry |
off | Enable noisy competitor price observations for firms |
--crash-rep-scoring |
off | THE_CRASH: score quotes by reputation/price instead of 1/price |
--dynamic-labor |
off | Re-sample CES labor each timestep (vs. fixed at t=0) |
--consumption-interval |
1 |
Run consumer inventory consumption every N timesteps |
--num-goods |
1 |
Number of goods in the simulation |
--fixed-consumer-quantity-per-good |
10.0 |
Quantity per good for FIXED consumers |
--listing-corpus |
— | LEMON_MARKET: path to pre-compiled listing corpus (eliminates seller LLM calls). See scripts/compile_listing_corpus.py |
--allow-listing-persistence |
off | LEMON_MARKET: carry unsold listings forward instead of discarding each step |
| Argument | Default | Description |
|---|---|---|
--reputation-initial |
0.8 |
Initial seller reputation R₀ (default 1.0 when no Sybil) |
--reputation-pseudo-count |
10 |
Rolling vote-window size N. reputation = upvotes in last N / N |
--sybil-rho-min |
0.3 |
Sybil rotation threshold: when R < rho_min, spawn new identity at --reputation-initial |
--no-buyer-rep |
off | Withhold seller reputation from buyer observations (ablation) |
--no-seller-ids |
off | Omit seller identifiers from buyer observations; listings get ephemeral per-round labels (ablation) |
--lemon-base-buyer |
off | Minimal buyer prompt with no transaction history (ablation) |
| Argument | Default | Description |
|---|---|---|
--llm |
llama3:8b |
Model name. Examples: gemini-2.5-flash, gpt-4o, meta-llama/llama-3.1-8b-instruct |
--buyer-llm |
— | LEMON_MARKET: model for buyer agents (falls back to --llm) |
--seller-llm |
— | LEMON_MARKET: model for honest sellers and Sybil principal (falls back to --llm) |
--stab-llm |
— | THE_CRASH: model name for Stabilizing Firms (e.g. a vLLM LoRA adapter alias like stab) |
--service |
vllm |
vllm or ollama for local models |
--buyer-service |
— | LEMON_MARKET: service for buyer agents (falls back to --service) |
--seller-service |
— | LEMON_MARKET: service for seller agents (falls back to --service) |
--port |
8009 |
Port for LLM service |
--buyer-port |
— | LEMON_MARKET: port for buyer LLM (falls back to --port) |
--seller-port |
— | LEMON_MARKET: port for seller LLM (falls back to --port) |
--gemini-backend |
auto | studio (API key) or vertex (Vertex AI). Auto-detects from env vars |
--openrouter-provider |
— | Preferred OpenRouter provider(s) (e.g. anthropic, Together) |
--buyer-openrouter-provider |
— | LEMON_MARKET: OpenRouter provider for buyer agents |
--seller-openrouter-provider |
— | LEMON_MARKET: OpenRouter provider for seller agents |
--prompt-algo |
io |
io (input-output) or cot (chain-of-thought) |
--history-len |
3 |
Timesteps of history sent in each firm prompt |
--best-n |
3 |
Best-N slab size for Stabilizing Firm prompts (0 to disable) |
--max-tokens |
1000 |
Maximum output tokens per LLM call |
--timeout |
30 |
LLM call timeout in seconds |
--use-parsing-agent |
off | Use a secondary LLM call to repair malformed JSON responses |
| Argument | Default | Description |
|---|---|---|
--name |
"" |
Run name (used as log directory label) |
--log-dir |
logs |
Base directory for output files |
--seed |
42 |
Random seed |
--wandb |
off | Enable Weights & Biases logging |
--no-diaries |
off | Disable strategic diary entries in agent prompts |
--log-firm-prompts |
off | Log firm prompt/response pairs to file |
--log-crash-firm-prompts |
off | THE_CRASH: append firm prompts to crash_agent_prompts.jsonl |
--log-buyer-prompts |
off | LEMON_MARKET: append buyer prompts to lemon_agent_prompts.jsonl |
--log-seller-prompts |
off | LEMON_MARKET: append seller/Sybil prompts to lemon_agent_prompts.jsonl |
--log-alignment-traces |
off | Log (state, prompt, response, outcome) tuples for SFT data collection |
--reward-type |
PROFIT |
PROFIT or REVENUE reward signal |
| Argument | Default | Description |
|---|---|---|
--shock-timestep |
— | Timestep at which to inject the shock |
--post-shock-unit-cost |
— | New unit cost after supply shock (THE_CRASH) |
--post-shock-sybil-cluster-size |
— | New Sybil cluster size after flood shock (LEMON_MARKET) |
All experiment scripts live in scripts/ and must be run from the project root. Use --list to preview runs without executing, and --skip-existing to resume partial sweeps.
5 LLM firms, 50 CES consumers, 365 timesteps. Sweeps stabilizing firm count and consumer discovery limit.
Full matrix (54 runs): baseline (k=0) over dlc ∈ {1,3,5} × seeds {8,16,64} + stabilizing-firm sweeps over dlc ∈ {1,3,5} × n_stab ∈ {1,2,3,4,5} × seeds {8,16,64}.
# Run all 54 cells sequentially
python scripts/exp1.py
# Parallel (keep workers low to respect API rate limits)
python scripts/exp1.py --workers 3
# Use a different model
python scripts/exp1.py --llm gemini-2.5-flash
# Local model via Ollama
python scripts/exp1.py --llm gemma3:4b --service ollama --port 11434
# Via OpenRouter with a specific provider
python scripts/exp1.py --llm anthropic/claude-sonnet-4-6 --openrouter-provider anthropic
# Filter: only dlc=3, n_stab=1 or 3, seed=8
python scripts/exp1.py --dlc 3 --n-stab 1 3 --seeds 8
# Resume a partial run
python scripts/exp1.py --skip-existing --workers 3
# Preview matching runs without launching (RECOMMENDED BEFORE LAUNCHING A JOB)
python scripts/exp1.py --dlc 3 --n-stab 1 3 --listFilter flags: --dlc, --n-stab, --seeds, --run (exact labels), --skip-existing
Logs: logs/exp1_<model>/
Runs the full Exp1 matrix for every dense open-weight model listed in [documentation/open_weights_models.json](documentation/open_weights_models.json) via OpenRouter. Add or remove models by editing that file — each entry is {"display_name": "...", "params_b": <float>, "slug": "<openrouter-slug>"}. The repo ships with a single example entry; populate it with the models you want to sweep. Use --models <substring> [<substring> …] to filter the loaded list.
# All models, 4 parallel workers
python scripts/exp1_eas_sweep.py --workers 4 --skip-existing
# Only dlc=3, k=3 cells (for the health-vs-size scatter)
python scripts/exp1_eas_sweep.py --dlc 3 --n-stab 3 --workers 4
# Subset of models by name substring
python scripts/exp1_eas_sweep.py --models llama-3.2-3b gemma-3-4b --workers 212 sellers (honest = 12 − K, Sybil = K), 12 LLM buyers, 50 timesteps. Sweeps Sybil saturation and reputation visibility.
Full matrix (24 runs): K ∈ {0,3,6,9} × rep_visible ∈ {True,False} × seeds {8,16,64}.
# All 24 cells
python scripts/exp2.py --seller-llm google/gemma-3-12b-it
# Parallel
python scripts/exp2.py --seller-llm google/gemma-3-12b-it --workers 3
# Split buyer and seller models
python scripts/exp2.py \
--seller-llm google/gemma-3-12b-it --seller-openrouter-provider Together \
--buyer-llm anthropic/claude-sonnet-4-6 --buyer-openrouter-provider anthropic
# Prompt logging
python scripts/exp2.py --log-buyer-prompts --log-seller-prompts
# Filter: only K=3 and K=6, rep hidden
python scripts/exp2.py --k 3 6 --rep-visible 0Filter flags: --k, --rep-visible (1=visible, 0=hidden), --seeds, --run, --skip-existing
Identical to Exp2 with --no-seller-ids hardwired, isolating whether buyers can detect lemons without cross-round seller tracking.
python scripts/exp2_2.py --llm gemini-2.5-flash --workers 3Sweeps buyer model capability against a fixed seller model.
python scripts/exp2_eas_sweep.py --seller-llm google/gemma-3-12b-it --workers 4 --skip-existingApplies mid-episode shocks to measure market resilience.
| Sub-experiment | Scenario | Shock | Timing |
|---|---|---|---|
| exp3a | THE_CRASH | Unit cost $1 → $10 | t = 25 |
| exp3b | LEMON_MARKET | Sybil cluster → 80% saturation | t = 15 |
Full matrix (36 runs): 18 crash (n_stab ∈ {1,3,5} × dlc ∈ {3,5} × seeds) + 18 lemon (k_initial ∈ {3,6,9} × rep_visible × seeds).
# All 36 runs
python scripts/exp3.py
# Only crash or lemon sub-experiment
python scripts/exp3.py --experiment crash
python scripts/exp3.py --experiment lemon
# Override the model under test
python scripts/exp3.py --test-llm anthropic/claude-sonnet-4-6 --openrouter-provider anthropicYou can also run shocks directly:
# Supply shock at t=25
python -m agent_bazaar.main \
--consumer-scenario THE_CRASH \
--firm-type LLM --num-firms 5 --num-consumers 50 \
--use-cost-pref-gen --overhead-costs 14 --max-timesteps 100 \
--shock-timestep 25 --post-shock-unit-cost 10.0 \
--llm gemini-3-flash-preview --seed 8
# Sybil flood at t=15
python -m agent_bazaar.main \
--consumer-scenario LEMON_MARKET \
--num-sellers 12 --num-buyers 12 \
--sybil-cluster-size 3 --reputation-initial 0.8 \
--max-timesteps 50 \
--shock-timestep 15 --post-shock-sybil-cluster-size 36 \
--llm gemini-3-flash-preview --seed 8Mirrors Exp1 but sweeps firm-side price discovery (--discovery-limit-firms) with consumer discovery held at dlc=3.
python scripts/exp5.py --workers 3
# Filter by DLF value
python scripts/exp5.py --dlf 3 --n-stab 1 2Tests demand-side heterogeneity with behavioral consumer personas (LOYAL, SMALL_BIZ, PRICE_HAWK, POPULAR, VARIETY) at dlc=5.
python scripts/exp6.py
# Parallel with a specific model
python scripts/exp6.py --llm gemini-2.5-flash --workers 3For large sweeps on a GPU cluster, start a single vLLM server with LoRA serving to route base model and adapter requests to one GPU:
# Start vLLM with LoRA adapters
python -m vllm.entrypoints.openai.api_server \
--model ./models/Qwen3.5-9B \
--enable-lora \
--lora-modules stab=./models/ai-bazaar-checkpoints/crash_stabilizer \
guardian=./models/ai-bazaar-checkpoints/lemon_guardian \
--port 8000 --gpu-memory-utilization 0.7
# Run exp1 against the base model, with Stabilizing Firms using the LoRA adapter
python scripts/exp1.py \
--llm ./models/Qwen3.5-9B \
--stab-llm stab \
--service vllm --port 8000 \
--workers 5 --skip-existingThe listing corpus (corpus/listing_corpus.json) eliminates seller LLM calls in lemon experiments, saving compute:
python -m agent_bazaar.main \
--consumer-scenario LEMON_MARKET \
--num-sellers 12 --num-buyers 12 \
--sybil-cluster-size 3 \
--listing-corpus corpus/listing_corpus.json \
--buyer-llm guardian --service vllm --port 8000 \
--max-timesteps 50 --seed 8
# Rebuild the corpus from existing logs
python scripts/compile_listing_corpus.py# Run the full test suite
pytest tests/ -v
# With coverage
pytest tests/ --cov=agent_bazaar --cov-report=html
# Individual test modules
pytest tests/test_bazaar_env.py -v
pytest tests/test_lemon_market.py -v
pytest tests/test_buyer_agent.py -vMost tests are self-contained. Tests that make live API calls (test_models.py, test_advanced_usage.py) require the relevant API keys.
Any model accessible via the following backends is supported:
| Backend | Flag | Examples |
|---|---|---|
| Google Gemini (AI Studio) | --llm gemini-2.5-flash |
gemini-2.5-flash, gemini-2.5-pro, gemini-3-flash-preview |
| Google Vertex AI | --llm gemini-2.5-flash --gemini-backend vertex |
Same model IDs |
| OpenAI | --llm gpt-4o |
gpt-4o, gpt-4o-mini, gpt-5.4 |
| Anthropic (via OpenRouter) | --llm anthropic/claude-sonnet-4-6 |
Any Anthropic model slug |
| OpenRouter | --llm org/model-name |
Any model on openrouter.ai |
| Ollama (local) | --service ollama --llm llama3.1:8b |
Any model in ollama list |
| vLLM (local) | --service vllm --llm hf/model-id |
Any HF model or LoRA alias |
The open-weight model list used by the EAS sweep scripts (scripts/exp1_eas_sweep.py, scripts/exp2_eas_sweep.py, scripts/exp3_open_weights_sweep.py) is loaded at runtime from [documentation/open_weights_models.json](documentation/open_weights_models.json). To add or remove models, edit that file — no code changes required.
If you use this framework, please cite the Agent Bazaar paper:
@misc{karten2026agentbazaar,
title = {The Agent Bazaar: Benchmarking Economic Alignment in High-Frequency Multi-Agent Ecosystems},
author = {Karten, Seth and Crow, Cameron and Jin, Chi},
year = {2026},
institution = {Princeton University},
note = {Under review}
}The simulator's consumer/firm scaffolding and LLM-call layer build on the LLM Economist:
@article{karten2025llm,
title = {LLM Economist: Large Population Models and Mechanism Design in Multi-Agent Generative Simulacra},
author = {Karten, Seth and Li, Wenzhe and Ding, Zihan and Kleiner, Samuel and Bai, Yu and Jin, Chi},
journal = {arXiv preprint arXiv:2507.15815},
year = {2025}
}MIT License — see LICENSE for details.

