Skip to content

docs(big-data): add DuckDB integration guide - #148

Merged
majinghe merged 1 commit into
rustfs:mainfrom
majinghe:docs/duckdb-integration
Sep 20, 2026
Merged

majinghe merged 1 commit into
rustfs:mainfrom
majinghe:docs/duckdb-integration

Conversation

@majinghe

Copy link
Copy Markdown
Collaborator

Summary

Adds a DuckDB guide under the Big Data category in all five locales (en, zh, de, fr, ja), wired into big-data/meta.json and the category landing pages.

  • Deploys RustFS with Docker Compose alongside the official duckdb/duckdb image (entrypoint: ["/duckdb"] — the image has no shell), with the rc-based create-bucket initializer.
  • Configures the httpfs S3 secret for the RustFS endpoint: ENDPOINT 'rustfs:9000', USE_SSL FALSE, URL_STYLE 'path' (path-style addressing).
  • Writes query results to the bucket with COPY ... TO 's3://my-bucket/duckdb-demo/events.parquet' (FORMAT PARQUET), reads them back with read_parquet, and verifies the object through rc ls and a Console screenshot (light theme, ≤300 KB; Chinese capture in zh, English in the others).
  • Also adds the missing PyIceberg entry to the big-data landing pages (overlooked when the PyIceberg guide landed).

Verification

Validated end to end on Ubuntu 24.04 against duckdb/duckdb:latest (v1.5.5) and rustfs/rustfs-x86-musl:v2.3.1:

  • read_parquet('s3://my-bucket/files/default/logs/...parquet') returned 200 rows from an existing object written by OpenObserve.
  • COPY (SELECT ... FROM range(1000)) TO 's3://my-bucket/duckdb-demo/events.parquet' succeeded; rc ls shows the 5.32 KiB object; reading it back returns rows: 1000, min_id: 0, max_id: 999.

npm run docs:check passes; npm run build passes (2535 pages); locale audit reports no findings for the new pages.

Add a DuckDB guide under the big data category in all five locales
(en, zh, de, fr, ja), and wire it into big-data meta.json and the
category landing pages.

The guide deploys RustFS with Docker Compose alongside the official
duckdb/duckdb image, configures the httpfs S3 secret for the RustFS
endpoint (path-style, plain HTTP), writes query results to the bucket
as Parquet with COPY TO, reads them back with read_parquet, and
verifies the objects through the rc CLI and the RustFS Console.

Also list PyIceberg in the big-data landing pages, which was missing
when the PyIceberg guide was added.

Verified end to end against duckdb/duckdb:latest (v1.5.5) and
rustfs/rustfs-x86-musl:v2.3.1: read 200 rows from an existing
Parquet object, wrote a 1000-row table to
s3://my-bucket/duckdb-demo/events.parquet, and read it back.
@vercel

vercel Bot commented Sep 20, 2026

Copy link
Copy Markdown

@majinghe is attempting to deploy a commit to the overtrue's projects Team on Vercel.

A member of the Team first needs to authorize it.

@majinghe
majinghe merged commit f281ddb into rustfs:main Sep 20, 2026
1 of 2 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant