Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
312 changes: 312 additions & 0 deletions Cargo.lock

Large diffs are not rendered by default.

43 changes: 43 additions & 0 deletions Cargo.toml
Original file line number Diff line number Diff line change
Expand Up @@ -71,10 +71,53 @@ cfb = "0.10"
# lopdf — the part we refuse to reimplement. Known to panic on hostile files;
# every call is wrapped in catch_unwind (see extract.rs).
pdf-extract = "0.12"
# OCR tier (ocr.rs): ocrs + rten — pure-Rust OCR. No tesseract or any system
# dependency (that would break the clean-VM stranger gate), no cloud OCR.
# Models (~12 MB total) download once from the ocrs project's canonical bucket
# into ~/.cache/ocrs — the same fetch-once-then-cache lane as the MiniLM
# encoder (which uses the hf-hub default cache). rten must track the exact
# version the ocrs release depends on, or Model types won't line up.
ocrs = "0.12"
rten = "0.24"
# image: PNG/JPEG decode for standalone image OCR and embedded PDF images,
# plus JPEG *encoding* in test fixtures. Decode-only features, pure Rust.
image = { version = "0.25", default-features = false, features = ["png", "jpeg"] }
# lopdf: pinned to the same version pdf-extract already compiles, so walking a
# scanned PDF's page/image tree reuses the one PDF parser in the tree.
lopdf = "0.42"
# flate2: capped FlateDecode inflation for embedded PDF images (ocr.rs). We
# deliberately do NOT use lopdf's decompressed_content() there — it inflates
# without any output bound, so a 100 KB stream declaring a 10x10 image can
# materialize gigabytes (decompression bomb). Already in the tree via zip.
flate2 = "1"
# ureq: blocking OCR-model download (extraction is synchronous code); already
# in the tree transitively via hf-hub.
ureq = "3"
reqwest = { version = "0.12", default-features = false, features = ["json", "rustls-tls", "multipart"] }
schemars = { version = "1", features = ["chrono04"] }
# Local-folder watching (folder_watch.rs): cross-platform FS events
# (FSEvents/inotify/ReadDirectoryChangesW) with a settle-debouncer so editor
# write bursts and mid-write partials fire once, after the file is quiescent.
notify = "6"
notify-debouncer-full = "0.3"

# OCR inference is numeric hot-loop work; unoptimized rten is ~50x slower,
# which makes dev-profile scanned-PDF ingestion (and the VERITY_OCR_E2E test)
# crawl. Optimizing JUST the rten family in dev keeps `cargo build` fast for
# our own code while making local OCR usable. Release builds are unaffected.
[profile.dev.package.rten]
opt-level = 3
[profile.dev.package.rten-base]
opt-level = 3
[profile.dev.package.rten-gemm]
opt-level = 3
[profile.dev.package.rten-imageproc]
opt-level = 3
[profile.dev.package.rten-simd]
opt-level = 3
[profile.dev.package.rten-tensor]
opt-level = 3
[profile.dev.package.rten-vecmath]
opt-level = 3
[profile.dev.package.ocrs]
opt-level = 3
13 changes: 13 additions & 0 deletions HONESTY.md
Original file line number Diff line number Diff line change
Expand Up @@ -98,6 +98,19 @@ source that has **never** heartbeated fails closed while the gate is on (never-s
indistinguishable from stalled). With the gate off — the default — assume a stalled
connector serves ACLs as stale as the stall is long, and monitor the heartbeats.

## OCR is local and printed-text-grade, not a document-AI service

Scanned PDFs and PNG/JPEG images are extracted by a fully local, pure-Rust OCR engine
([ocrs](https://github.com/robertknight/ocrs) on rten; ~12 MB of models fetched once into
`~/.cache/ocrs`, no cloud calls, no system dependencies). It reads printed type well and is
**best-effort beyond that**: expect it to miss handwriting, low-resolution scans, stylized
layouts, and non-Latin scripts. Every extraction receipt discloses the method — `pdf-ocr`
(with `pages_ocred`, capped at 50 pages) or `image-ocr` — so OCR-derived text is never
passed off as a document's own text layer, and a file where OCR finds nothing lands
metadata-only with a typed reason rather than silently indexing empty. Encrypted PDFs are
still declined outright. Unsupported embedded encodings (CCITT/JBIG2/JPEG 2000) are
skipped, not guessed at.

## Derived-visibility on agent writes: lineage is client-DECLARED, not inferred

`remember` (`POST /v1/episodes`) now enforces the SPEC §2 intersection
Expand Down
60 changes: 43 additions & 17 deletions crates/verity-mcp/src/main.rs
Original file line number Diff line number Diff line change
Expand Up @@ -246,8 +246,9 @@ struct IngestTextParams {
struct IngestFileParams {
/// Scope handle from memory_open_scope.
scope_handle: String,
/// Path to a LOCAL UTF-8 text file. Allowed extensions: .txt, .md,
/// .json, .csv, .html. Maximum size: 512 KB.
/// Path to a LOCAL file. Allowed extensions: .txt, .md, .json, .csv,
/// .html (UTF-8 text) and .png, .jpg, .jpeg (server-side local OCR).
/// Maximum size: 512 KB.
path: String,
/// Entity tags, e.g. ["account:acme-corp"]. Must be inside the scope's
/// entity_scope; omit to inherit the whole scope.
Expand Down Expand Up @@ -375,22 +376,20 @@ impl VerityMcp {
}

/// Multipart POST /v1/files: fields `scope_handle`, `entities`
/// (comma-separated, only when tags were given), and `file`.
/// (comma-separated, only when tags were given), and `file`. The part is
/// caller-built: text content for the UTF-8 lane, raw bytes + image mime
/// for the OCR lane (the server extracts by magic either way).
async fn post_file(
&self,
scope_handle: String,
entities: Option<Vec<String>>,
file_name: String,
content: String,
part: reqwest::multipart::Part,
) -> Result<CallToolResult, ErrorData> {
let mut form = reqwest::multipart::Form::new().text("scope_handle", scope_handle);
if let Some(entities) = entities.filter(|e| !e.is_empty()) {
form = form.text("entities", entities.join(","));
}
form = form.part(
"file",
reqwest::multipart::Part::text(content).file_name(file_name),
);
form = form.part("file", part);
self.proxy(self.http.post(self.endpoint("/v1/files")).multipart(form))
.await
}
Expand All @@ -404,8 +403,12 @@ fn tool_error(msg: impl Into<String>) -> Result<CallToolResult, ErrorData> {

// ---------- local helpers for the ingest tools ----------

/// Extensions memory_ingest_file accepts (UTF-8 text-like content only).
/// Extensions memory_ingest_file accepts as UTF-8 text.
const INGEST_FILE_EXTENSIONS: [&str; 5] = ["txt", "md", "json", "csv", "html"];
/// Extensions memory_ingest_file accepts as raw image bytes: the server's
/// local OCR tier (extract.rs + ocr.rs) extracts printed text best-effort,
/// disclosed as method "image-ocr" on the receipt.
const INGEST_IMAGE_EXTENSIONS: [&str; 3] = ["png", "jpg", "jpeg"];
/// memory_ingest_file size cap.
const MAX_FILE_BYTES: u64 = 512 * 1024;
/// memory_ingest_url download cap.
Expand Down Expand Up @@ -849,7 +852,7 @@ impl VerityMcp {

#[tool(
name = "memory_ingest_file",
description = "Read a LOCAL text file and ingest its contents into shared memory, so it becomes searchable via memory_recall. Use when the knowledge lives in a file on this machine rather than in-context. Accepts UTF-8 .txt/.md/.json/.csv/.html up to 512 KB; anything else is rejected with an error."
description = "Read a LOCAL file and ingest its contents into shared memory, so it becomes searchable via memory_recall. Use when the knowledge lives in a file on this machine rather than in-context. Accepts UTF-8 text (.txt/.md/.json/.csv/.html) and images (.png/.jpg/.jpeg — printed text is extracted by the server's local OCR, best-effort, disclosed as method image-ocr) up to 512 KB; anything else is rejected with an error."
)]
async fn memory_ingest_file(
&self,
Expand All @@ -861,9 +864,10 @@ impl VerityMcp {
.and_then(|e| e.to_str())
.map(str::to_ascii_lowercase)
.unwrap_or_default();
if !INGEST_FILE_EXTENSIONS.contains(&ext.as_str()) {
let is_image = INGEST_IMAGE_EXTENSIONS.contains(&ext.as_str());
if !is_image && !INGEST_FILE_EXTENSIONS.contains(&ext.as_str()) {
return tool_error(format!(
"unsupported file type {:?}: memory_ingest_file accepts only UTF-8 text files with extension .txt, .md, .json, .csv, or .html",
"unsupported file type {:?}: memory_ingest_file accepts UTF-8 text files (.txt, .md, .json, .csv, .html) and images (.png, .jpg, .jpeg)",
p.path
));
}
Expand All @@ -885,6 +889,28 @@ impl VerityMcp {
Ok(bytes) => bytes,
Err(e) => return tool_error(format!("cannot read file {:?}: {e}", p.path)),
};
if is_image {
// Raw bytes to the server's OCR lane; the server sniffs magic and
// returns a typed, disclosed failure if OCR finds nothing.
let file_name = path
.file_name()
.and_then(|n| n.to_str())
.unwrap_or("file.png")
.to_owned();
let mime = if ext == "png" {
"image/png"
} else {
"image/jpeg"
};
let part = match reqwest::multipart::Part::bytes(bytes)
.file_name(file_name)
.mime_str(mime)
{
Ok(part) => part,
Err(e) => return tool_error(format!("building upload part: {e}")),
};
return self.post_file(p.scope_handle, p.entities, part).await;
}
let content = match String::from_utf8(bytes) {
Ok(content) => content,
Err(_) => {
Expand All @@ -899,8 +925,8 @@ impl VerityMcp {
.and_then(|n| n.to_str())
.unwrap_or("file.txt")
.to_owned();
self.post_file(p.scope_handle, p.entities, file_name, content)
.await
let part = reqwest::multipart::Part::text(content).file_name(file_name);
self.post_file(p.scope_handle, p.entities, part).await
}

#[tool(
Expand Down Expand Up @@ -980,8 +1006,8 @@ impl VerityMcp {
return tool_error(format!("no textual content extracted from {url}"));
}
let file_name = file_name_from_url(&url);
self.post_file(p.scope_handle, p.entities, file_name, content)
.await
let part = reqwest::multipart::Part::text(content).file_name(file_name);
self.post_file(p.scope_handle, p.entities, part).await
}

#[tool(
Expand Down
25 changes: 23 additions & 2 deletions crates/verity-server/Cargo.toml
Original file line number Diff line number Diff line change
Expand Up @@ -42,16 +42,37 @@ moka = { workspace = true }
# Media object-store seam (task 47): S3/MinIO blob tier behind object_store.
object_store = { workspace = true }
bytes = "1"
# Tier-1 file extraction (extract.rs): PDF/PPTX/XLS(X) → text, deterministic,
# no OCR. Dep choices justified at the workspace root Cargo.toml.
# Tier-1 file extraction (extract.rs): PDF/PPTX/XLS(X)/DOC(X) text layers,
# deterministic. Dep choices justified at the workspace root Cargo.toml.
calamine = { workspace = true }
zip = { workspace = true }
cfb = { workspace = true }
quick-xml = { workspace = true }
pdf-extract = { workspace = true }
# OCR tier (ocr.rs): image decode + scanned-PDF page walking are always built
# (typed failures either way); the ocrs/rten engine itself sits behind the
# default-ON `ocr` feature so `--no-default-features` can shave its compile
# cost — extraction then fails typed ("built without the 'ocr' feature").
image = { workspace = true }
lopdf = { workspace = true }
# Bomb-guarded FlateDecode for embedded PDF images: lopdf's own
# decompressed_content() has no output cap, so ocr.rs inflates manually with
# flate2 + Read::take. Non-optional — the decode path compiles either way.
flate2 = { workspace = true }
ocrs = { workspace = true, optional = true }
rten = { workspace = true, optional = true }
ureq = { workspace = true, optional = true }
# SSE subscriptions (task 21): poll-loop event streams.
async-stream = "0.3"
futures-util = { version = "0.3", default-features = false, features = ["alloc"] }
# Local-folder watching (folder_watch.rs): server-side FS watcher + debounce.
notify = { workspace = true }
notify-debouncer-full = { workspace = true }

[features]
default = ["ocr"]
# Local OCR engine (ocrs + rten). ON by default — sovereignty-first, still
# pure Rust, still zero system deps. Building with --no-default-features
# drops the engine (and its compile time); scanned-PDF/image extraction then
# returns the typed "built without the 'ocr' feature" failure instead.
ocr = ["dep:ocrs", "dep:rten", "dep:ureq"]
Loading
Loading