Skip to content
brecabralPublic

About

Run an LLM over a whole batch of inputs — reliably. sluice reads a file (one prompt per row), sends each item to an LLM under controlled concurrency with retries, rate limiting and timeouts, and writes one result per item plus a summary. Everything is driven by a single sluice.yaml.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Latest commit

 

History

15 Commits

Folders and files

Repository files navigation

sluice

Run an LLM over a whole batch of inputs — reliably. sluice reads a file (one prompt per row), sends each item to an LLM under controlled concurrency with retries, rate limiting and timeouts, and writes one result per item plus a summary. Everything is driven by a single sluice.yaml.

Build

Requires Go 1.26+.

go build -o sluice ./cmd/cli

Quickstart

  1. Run a model locally with Ollama:

    ollama pull llama3.2
  2. Write a sluice.yaml (see below), then run it from the folder that contains it:

    ./sluice

    It reads the input file, processes every row, and writes the output file. While running (in a terminal) it shows live progress and, at the end, a summary:

    processing 10/10 (0 failed)
    10 items · 10 done · 0 failed · 7.3 items/s (1.37s)
    output written to out.json
    

Configuration (sluice.yaml)

provider:
  kind: ollama            # test | ollama | generic
  options:
    model: llama3.2

task:
  instruction: Classify the sentiment of the review.
  # Omit `output` for free-text mode (the raw model reply is stored as `output`).
  # Declare `output` fields for structured output — each field optionally an enum
  # via `values` (validated; invalid replies are retried):
  output:
    - name: sentiment
      description: overall sentiment
      values: [positive, negative, neutral]
    - name: reason
      description: short justification

input:
  path: examples/reviews.csv   # relative to the config file's directory
  options:
    column: comment            # which CSV column to read

output:
  path: out.json               # kind inferred from the extension

worker:
  concurrency: 4               # workers processing in parallel
  max_attempts: 5              # attempts per item (retry on error/invalid output)
  rate_limit: 0                # provider calls per second (0 = unlimited)
  timeout: 60s                 # per-request timeout ("" = none)
  backoff: 500ms               # base retry backoff, grows exponentially, 4x on 429

Providers

  • test — no-op, returns nothing. Useful to dry-run the pipeline offline.
  • ollama — a local Ollama server. options.model (required); options.url defaults to http://localhost:11434/api/generate.
  • generic — any HTTP endpoint that accepts {"prompt": "..."} and replies {"response": "..."}. options.url (required); options.api_key_env names the environment variable holding the API key (the key itself never goes in the YAML) and is sent as Authorization: Bearer <key>.

Choosing the config file

By default sluice reads sluice.yaml from the current directory (where you run it — not where the binary lives). Point at another config with -config (shorthand -c):

sluice -config path/to/sluice.yaml

Relative input.path / output.path inside a config resolve against that config's directory, so a folder with its own sluice.yaml + data is self-contained and runs from anywhere.

Examples

The examples/ folder has ready-to-run configs (Ollama) covering different uses:

Config Use case
examples/sentiment.yaml classify review sentiment (enum + reason)
examples/summarize.yaml summarize articles (free-text output)
examples/extract.yaml extract category + priority + summary
examples/triage.yaml assign a support priority (enum)

Run one against your local Ollama:

just run-example sentiment
# writes examples/sentiment.out.json

Output

out.json holds a summary and one entry per item:

{
  "summary": { "Total": 10, "Failed": 0, "ElapsedSecs": 1.37, "ItemsPerSec": 7.3 },
  "results": [
    { "input": "great product", "fields": { "sentiment": "positive", "reason": "..." }, "status": "done", "attempts": 1 }
  ]
}

In free-text mode each result has output (the raw reply) instead of fields.

Press Ctrl-C mid-run to stop gracefully: in-flight calls are cancelled and the partial results gathered so far are still written.

About

Run an LLM over a whole batch of inputs — reliably. sluice reads a file (one prompt per row), sends each item to an LLM under controlled concurrency with retries, rate limiting and timeouts, and writes one result per item plus a summary. Everything is driven by a single sluice.yaml.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Contributors

Languages