diff --git a/content/de/developer/integration/ai/images/rustfs-vllm-models.png b/content/de/developer/integration/ai/images/rustfs-vllm-models.png new file mode 100644 index 00000000..baa18dc1 Binary files /dev/null and b/content/de/developer/integration/ai/images/rustfs-vllm-models.png differ diff --git a/content/de/developer/integration/ai/index.md b/content/de/developer/integration/ai/index.md index bad31e8a..4d60d869 100644 --- a/content/de/developer/integration/ai/index.md +++ b/content/de/developer/integration/ai/index.md @@ -8,5 +8,6 @@ Nutzen Sie **RustFS** als Objektspeicher-Layer für KI- und Machine-Learning-Pla ## Plattformen - [Ray](./ray.md) +- [vLLM](./vllm.md) Speichern Sie Trainingsdaten und Checkpoints in dedizierten Buckets und beschränken Sie die Anmeldeinformationen auf die erforderlichen Bucket-Operationen. diff --git a/content/de/developer/integration/ai/meta.json b/content/de/developer/integration/ai/meta.json index 573fc520..b30f9e0f 100644 --- a/content/de/developer/integration/ai/meta.json +++ b/content/de/developer/integration/ai/meta.json @@ -1,6 +1,7 @@ { "title": "AI", "pages": [ - "ray" + "ray", + "vllm" ] } diff --git a/content/de/developer/integration/ai/vllm.md b/content/de/developer/integration/ai/vllm.md new file mode 100644 index 00000000..c471096c --- /dev/null +++ b/content/de/developer/integration/ai/vllm.md @@ -0,0 +1,154 @@ +--- +title: "vLLM" +description: "Serve LLM inference with vLLM loading model weights stored in RustFS." +--- + +This guide connects [vLLM](https://github.com/vllm-project/vllm) — the high-throughput LLM inference engine — to **RustFS** as its model-weight store. You will upload a model into a RustFS bucket, expose the bucket to the vLLM host through an rclone mount, and serve the model with the OpenAI-compatible API. The workflow was verified with `vllm/vllm-openai-cpu` (vLLM 0.30.0) serving `facebook/opt-125m` from a RustFS bucket backed by `rustfs/rustfs-x86-musl:v2.3.1`, on a CPU-only host. + +You need Docker and an rclone binary on the host that runs vLLM. This deployment is intended for local integration testing, not production. + +## Architecture + +```mermaid +flowchart LR + Upload["rclone copy"] -->|"weights"| RustFS["RustFS :9000"] + RustFS -->|"rclone mount"| Mount["/mnt/vllm-models"] + Mount -->|"weight load"| vLLM["vLLM :8000"] + Client["OpenAI SDK / curl"] -->|"completions"| vLLM +``` + +The bucket is the single copy of the model. Hosts that serve the model mount the bucket read-only, so every node pulls weights from RustFS and no local model store exists to drift. + +:::note[Why an rclone mount] + +vLLM 0.30 loads `s3://` model paths through the RunAI model streamer, whose ranged reads currently fail against custom S3 endpoints such as RustFS (the loader errors with `File access error` on any non-zero offset). Mounting the bucket as a filesystem is the verified way to keep the weights in RustFS while vLLM reads them as local files. + +::: + +## 1. Upload the model to RustFS + +Create the bucket and copy model weights into it, replacing all connection placeholders: + +```ini title="rclone.conf" +[rustfs] +type = s3 +provider = Other +access_key_id = +secret_access_key = +endpoint = http://:9000 +region = us-east-1 +``` + +```bash +rc mb rustfs/vllm-models +rclone copy ./opt-125m rustfs:vllm-models/opt-125m --transfers 4 +``` + +Any Hugging Face layout works — `config.json`, the tokenizer files, and the weight files (`model.safetensors` or `pytorch_model.bin`). Keep one model per prefix so several models can share the bucket. + +## 2. Mount the bucket on the vLLM host + +On the machine that runs vLLM, mount the bucket read-only for clients with `--allow-other`: + +```bash +mkdir -p /mnt/vllm-models +rclone mount rustfs:vllm-models /mnt/vllm-models \ + --allow-other --daemon +ls /mnt/vllm-models/opt-125m/ +``` + +```text +config.json merges.txt model.safetensors tokenizer.json vocab.json +``` + +## 3. Run vLLM + +Start the CPU image against the mounted weights: + +```bash +docker run -d --name vllm -p 8000:8000 --shm-size=2g \ + -v /mnt/vllm-models:/models:ro \ + vllm/vllm-openai-cpu:latest \ + --model /models/opt-125m --served-model-name opt-125m \ + --dtype float32 --max-model-len 256 --gpu-memory-utilization 0.15 +``` + +vLLM reads the weights through the mount — the container stays stateless and the model lives in RustFS. Wait for the server to come up: + +```bash +curl -s http://localhost:8000/v1/models | head -c 200 +``` + +```text +{"object":"list","data":[{"id":"opt-125m","object":"model","created":...,"root":"/models/opt-125m",...}]} +``` + +`--gpu-memory-utilization` controls the fraction of RAM reserved for the KV cache on the CPU backend; lower it on small hosts. `--dtype float32` matches what the CPU attention kernels support for this model. + +## 4. Run inference + +Send an OpenAI-compatible completion request: + +```bash +curl -s http://localhost:8000/v1/completions \ + -H "Content-Type: application/json" \ + -d '{"model": "opt-125m", "prompt": "RustFS is", "max_tokens": 12, "temperature": 0}' +``` + +```json +{"id":"cmpl-...","object":"text_completion","model":"opt-125m", + "choices":[{"index":0,"text":" a great tool for building your own server. It's a", + "finish_reason":"length",...}]} +``` + +The request is standard OpenAI schema, so the Python client works unchanged: + +```python +from openai import OpenAI + +client = OpenAI(base_url="http://localhost:8000/v1", api_key="EMPTY") +print(client.completions.create( + model="opt-125m", prompt="RustFS is", max_tokens=12, temperature=0, +).choices[0].text) +``` + +![vLLM model weights stored in the RustFS Console](./images/rustfs-vllm-models.png) + +## 5. Stop or reset + +To tear down the demo while keeping the bucket objects: + +```bash +docker rm -f vllm +fusermount -u /mnt/vllm-models +``` + +To delete the stored model: + +```bash +rclone purge rustfs:vllm-models +``` + +## Troubleshooting + +### `Cannot find any model weights with /models/...` + +The mount had a stale directory cache or the weight files never made it to the bucket. Re-run `rclone copy` and confirm the files through the mount with `ls` before starting vLLM. A short `--dir-cache-time` (for example `10s`) helps while you iterate. + +### `Unsupported CPU attention configuration: head_dim=...` + +vLLM's CPU kernels support a fixed set of head dimensions. Tiny test models such as `hf-internal-testing/tiny-random-*` use exotic shapes that fail at request time — use a real small model such as `facebook/opt-125m`. + +### `Insufficient space in /dev/shm` + +vLLM's CPU engine exchanges tensors through shared memory. Run the container with `--shm-size=2g` (or `--ipc=host`). + +### Server exits with `Available memory on node 0 ... is less than desired CPU memory utilization` + +The default KV-cache reservation is 90% of system RAM. Lower it with `--gpu-memory-utilization 0.15` (the flag applies to the CPU backend as a memory fraction despite its name). + +## Next steps + +- Review [S3 compatibility notes](/administration/protocols/s3) before adopting additional serving setups. +- Create dedicated production credentials with [Access Key Management](/security-compliance/iam/access-token). +- Follow the [vLLM documentation](https://docs.vllm.ai/en/latest/serving/openai_compatible_server.html) for chat templates, tensor parallelism, and quantized weights on top of the same bucket-backed model store. diff --git a/content/de/developer/integration/big-data/airflow.md b/content/de/developer/integration/big-data/airflow.md new file mode 100644 index 00000000..7f55e219 --- /dev/null +++ b/content/de/developer/integration/big-data/airflow.md @@ -0,0 +1,170 @@ +--- +title: "Airflow" +description: "Move data between Airflow DAGs and RustFS with the Amazon S3 provider." +--- + +This guide connects [Apache Airflow](https://github.com/apache/airflow) — the workflow orchestration platform — to **RustFS** through the Amazon S3 provider's hooks, operators, and sensors. You will register a custom-endpoint connection, run a DAG that writes an object to a RustFS bucket, waits for a key with `S3KeySensor`, and reads the object back with `S3Hook`. The workflow was verified with `apache/airflow:3.3.2` (standalone, SequentialExecutor) and the `apache-airflow-providers-amazon` provider against `rustfs/rustfs-x86-musl:v2.3.1`. + +You need Docker, or an existing Airflow installation. This deployment is intended for local integration testing, not production. + +## Architecture + +```mermaid +flowchart LR + Scheduler["Airflow scheduler"] -->|"tasks"| Hook["S3Hook / operators"] + Hook -->|"S3 API"| RustFS["RustFS :9000"] + Sensor["S3KeySensor"] -->|"poll key"| RustFS +``` + +Every S3 interaction inside a DAG goes through the provider's S3 client, pointed at RustFS by the connection's `endpoint_url`. Operators, sensors, and hooks share the same connection object. + +## 1. Run Airflow + +Start a standalone instance with examples disabled, and create the demo bucket: + +```bash +docker run -d --name airflow --network oo-rustfs_default -p 8080:8080 \ + -e AIRFLOW__CORE__LOAD_EXAMPLES=False \ + -v "$PWD/dags":/opt/airflow/dags \ + apache/airflow:3.3.2 standalone + +rc mb rustfs/airflow-demo +``` + +The image ships with all providers preinstalled, including `apache-airflow-providers-amazon`. + +## 2. Register the RustFS connection + +The S3 provider reads its endpoint from the connection's extra field. Replace all connection placeholders: + +```bash +docker exec airflow airflow connections add rustfs \ + --conn-type aws \ + --conn-extra '{"endpoint_url": "http://:9000", "region_name": "us-east-1", "aws_access_key_id": "", "aws_secret_access_key": ""}' +``` + +The keys `aws_access_key_id` and `aws_secret_access_key` inside `--conn-extra` supply credentials; `endpoint_url` redirects the boto3 client from AWS to RustFS. + +## 3. Write the DAG + +The DAG writes an object with an operator, waits for the key with a sensor, and reads it back with the hook: + +```python title="rustfs_demo.py" +import datetime + +from airflow.providers.amazon.aws.hooks.s3 import S3Hook +from airflow.providers.amazon.aws.operators.s3 import S3CreateObjectOperator +from airflow.providers.amazon.aws.sensors.s3 import S3KeySensor +from airflow.sdk import dag, task + +@dag( + schedule=None, + start_date=datetime.datetime(2026, 1, 1), + catchup=False, + tags=["rustfs"], +) +def rustfs_demo(): + create = S3CreateObjectOperator( + task_id="write_object", + s3_bucket="airflow-demo", + s3_key="dags/airflow-put.txt", + data="written by airflow to rustfs", + aws_conn_id="rustfs", + replace=True, + ) + + wait = S3KeySensor( + task_id="wait_for_object", + bucket_key="dags/airflow-put.txt", + bucket_name="airflow-demo", + aws_conn_id="rustfs", + timeout=120, + poke_interval=10, + mode="reschedule", + ) + + @task + def read_object(): + hook = S3Hook(aws_conn_id="rustfs") + body = hook.read_key(key="dags/airflow-put.txt", bucket_name="airflow-demo") + print("read back:", body) + assert body == "written by airflow to rustfs" + + create >> [wait, read_object()] + +rustfs_demo() +``` + +Note the import paths: `S3CreateObjectOperator` lives in the `operators` module while `S3KeySensor` lives in the `sensors` module — importing both from one place fails. + +## 4. Unpause and trigger + +New DAGs start paused, and a trigger fired while paused stays queued forever. Unpause first, then trigger: + +```bash +docker exec airflow airflow dags unpause rustfs_demo +docker exec airflow airflow dags trigger rustfs_demo +``` + +Watch the run finish: + +```bash +docker exec airflow airflow dags list-runs rustfs_demo | head -3 +``` + +```text +dag_id run_id state +rustfs_demo manual__2026-09-29T13:15:57.332688+00:00 success +``` + +All three tasks succeed: `write_object`, `wait_for_object`, and `read_object`. + +## 5. Verify objects in RustFS + +List the bucket prefix: + +```bash +rc ls rustfs/airflow-demo/ -r +rc cat rustfs/airflow-demo/dags/airflow-put.txt +``` + +```text +[2026-09-29 13:16:01] 28 B dags/airflow-put.txt +written by airflow to rustfs +``` + +![Airflow object stored in the RustFS Console](./images/rustfs-airflow-object.png) + +## 6. Stop or reset + +To tear down the demo while keeping the bucket objects: + +```bash +docker rm -f airflow +``` + +To delete the stored data: + +```bash +rc rm rustfs/airflow-demo/ --recursive --force +``` + +## Troubleshooting + +### Dag runs stay `queued` after triggering + +The DAG is paused. New DAGs are paused by default in Airflow 3, and runs triggered in that state never execute. Run `airflow dags unpause rustfs_demo`; queued runs then start on their own. + +### `cannot import name 'S3KeySensor' from 'airflow.providers.amazon.aws.operators.s3'` + +The sensor lives in a separate module: `from airflow.providers.amazon.aws.sensors.s3 import S3KeySensor`. + +### Tasks fail with connection errors + +The `endpoint_url` must be reachable from the Airflow container — use the Docker network hostname for RustFS, not `localhost`. Airflow 3 serves its health endpoint under `/api/v2/monitor/health` if you need to check component status. + +## Next steps + +- Review [S3 compatibility notes](/administration/protocols/s3) before adopting additional provider hooks. +- Create dedicated production credentials with [Access Key Management](/security-compliance/iam/access-token). +- Follow the [Amazon provider documentation](https://airflow.apache.org/docs/apache-airflow-providers-amazon/stable/index.html) for transfer operators such as `S3ToLocalFilesystemOperator` and `LocalFilesystemToS3Operator`. diff --git a/content/de/developer/integration/big-data/delta-lake.md b/content/de/developer/integration/big-data/delta-lake.md new file mode 100644 index 00000000..b1bfae56 --- /dev/null +++ b/content/de/developer/integration/big-data/delta-lake.md @@ -0,0 +1,125 @@ +--- +title: "Delta Lake" +description: "Write and read Delta tables on RustFS with delta-rs." +--- + +This guide connects [Delta Lake](https://github.com/delta-io/delta) — the open-source lakehouse table format — to **RustFS** through delta-rs, the Rust-native Delta implementation. You will write a Delta table to a RustFS bucket from Python, read it back with ACID transaction history, and confirm the `_delta_log` and Parquet files in the bucket. The workflow was verified with the `deltalake` Python package (delta-rs) and pandas against `rustfs/rustfs-x86-musl:v2.3.1`. + +You need Python 3.9 or newer. This deployment is intended for local integration testing, not production. + +## Architecture + +```mermaid +flowchart LR + DF["pandas DataFrame"] -->|"write_deltalake"| deltaRS["delta-rs"] + deltaRS -->|"Parquet + _delta_log"| RustFS["RustFS :9000"] + Query["DeltaTable"] -->|"read / time travel"| RustFS +``` + +delta-rs stores each table as Parquet files plus a transaction log (`_delta_log/`). All I/O goes through the `object_store` crate, configured with the same AWS environment variables as other S3 clients. + +## 1. Install the client + +```bash +pip install deltalake pandas pyarrow +``` + +`pyarrow` is required to convert pandas frames into Delta-compatible record batches. + +## 2. Write a Delta table + +Create the bucket and write a table, replacing all connection placeholders. `AWS_S3_ALLOW_UNSAFE_RENAME` is needed because RustFS does not provide copy-if-not-exists, which delta-rs otherwise uses for commit conflicts: + +```python title="delta_s3.py" +import pandas as pd +from deltalake import DeltaTable, write_deltalake + +storage_options = { + "AWS_ENDPOINT_URL": "http://:9000", + "AWS_ACCESS_KEY_ID": "", + "AWS_SECRET_ACCESS_KEY": "", + "AWS_REGION": "us-east-1", + "AWS_ALLOW_HTTP": "true", + "AWS_S3_ALLOW_UNSAFE_RENAME": "true", +} + +table = "s3:///events" +df = pd.DataFrame({"id": [1, 2, 3], "name": ["alpha", "beta", "gamma"]}) +write_deltalake(table, df, storage_options=storage_options) +print("written:", df.shape[0], "rows") +``` + +```text +written: 3 rows +``` + +The `table` URI uses the standard `s3://bucket/prefix` form; the endpoint and credentials come from `storage_options`. + +## 3. Read the table back + +```python title="delta_read.py" +from deltalake import DeltaTable + +back = DeltaTable("s3:///events", storage_options=storage_options).to_pandas() +print("read back:", back.shape[0], "rows") +print(back.sort_values("id").to_string(index=False)) +print("version:", DeltaTable("s3:///events", storage_options=storage_options).version()) +``` + +```text +read back: 3 rows + id name + 1 alpha + 2 beta + 3 gamma +version: 0 +``` + +Because the version is tracked in the transaction log, the same table supports time travel with `DeltaTable(..., version=N)` and appends that bump the version. + +## 4. Verify objects in RustFS + +List the table prefix: + +```bash +rc ls rustfs// -r +``` + +The first commit created the transaction log and one Parquet file: + +```text +events/_delta_log/00000000000000000000.json +events/part-00000-3859855e-45e4-4ae5-94ff-2d8eab5e7ebb-c000.snappy.parquet +``` + +Every new write adds a `NNNNNNNNNNNNNNNNNNNN.json` log entry and Parquet parts; readers replay the log to get a consistent snapshot. + +![Delta table files stored in the RustFS Console](./images/rustfs-delta-table.png) + +## 5. Stop or reset + +delta-rs holds no state of its own. To delete the table: + +```bash +rc rm rustfs//events/ --recursive --force +``` + +## Troubleshooting + +### `Import pyarrow failed` when writing a pandas DataFrame + +`write_deltalake` converts frames through Arrow. Install `pyarrow` alongside `deltalake` and `pandas`. + +### `Generic DeltaTable error: commit conflict` or rename errors on commit + +delta-rs commits by copying and renaming temporary objects, which requires atomic rename on the backend. For S3-compatible stores without copy-if-not-exists, set `AWS_S3_ALLOW_UNSAFE_RENAME: "true"` in `storage_options` — acceptable for a single writer, not for concurrent writers. + +### `Unknown lengthy error: AWS connectivity or endpoint errors` + +Confirm `AWS_ENDPOINT_URL` includes the scheme and that `AWS_ALLOW_HTTP` is `"true"` for plain-HTTP endpoints; without it the S3 client only speaks HTTPS. + +## Next steps + +- Review [S3 compatibility notes](/administration/protocols/s3) before adopting additional Delta clients. +- Create dedicated production credentials with [Access Key Management](/security-compliance/iam/access-token). +- Follow the [delta-rs usage documentation](https://delta-io.github.io/delta-rs/usage/writing/writing-to-s3/) for concurrent-writer setups with DynamoDB-backed commit coordination. diff --git a/content/de/developer/integration/big-data/images/rustfs-airflow-object.png b/content/de/developer/integration/big-data/images/rustfs-airflow-object.png new file mode 100644 index 00000000..9177d249 Binary files /dev/null and b/content/de/developer/integration/big-data/images/rustfs-airflow-object.png differ diff --git a/content/de/developer/integration/big-data/images/rustfs-delta-table.png b/content/de/developer/integration/big-data/images/rustfs-delta-table.png new file mode 100644 index 00000000..093a5cf1 Binary files /dev/null and b/content/de/developer/integration/big-data/images/rustfs-delta-table.png differ diff --git a/content/de/developer/integration/big-data/images/rustfs-kafka-sink.png b/content/de/developer/integration/big-data/images/rustfs-kafka-sink.png new file mode 100644 index 00000000..dd4b11b5 Binary files /dev/null and b/content/de/developer/integration/big-data/images/rustfs-kafka-sink.png differ diff --git a/content/de/developer/integration/big-data/index.md b/content/de/developer/integration/big-data/index.md index 5931816f..72ebcc11 100644 --- a/content/de/developer/integration/big-data/index.md +++ b/content/de/developer/integration/big-data/index.md @@ -8,6 +8,7 @@ Use **RustFS** as the object storage layer for data analytics systems that suppo ## Systems - [ClickHouse](./clickhouse.md) +- [Airflow](./airflow.md) - [Hudi](./hudi.md) - [Iceberg](./iceberg.md) - [PyIceberg](./pyiceberg.md) @@ -16,8 +17,10 @@ Use **RustFS** as the object storage layer for data analytics systems that suppo - [OpenDAL](./opendal.md) - [DuckDB](./duckdb.md) - [Doris](./doris.md) +- [Delta Lake](./delta-lake.md) - [lakeFS](./lakefs.md) - [InfluxDB](./influxdb.md) +- [Kafka](./kafka.md) - [Spark](./spark.md) - [Flink](./flink.md) - [Trino](./trino.md) diff --git a/content/de/developer/integration/big-data/kafka.md b/content/de/developer/integration/big-data/kafka.md new file mode 100644 index 00000000..ff1e92f9 --- /dev/null +++ b/content/de/developer/integration/big-data/kafka.md @@ -0,0 +1,178 @@ +--- +title: "Kafka" +description: "Offload Kafka topic data to RustFS with the Kafka Connect S3 sink connector." +--- + +This guide connects [Apache Kafka](https://github.com/apache/kafka) — the distributed event streaming platform — to **RustFS** through the Kafka Connect S3 sink connector. You will run a KRaft broker and a Connect worker, deploy the S3 sink for a topic, and produce records that land as objects in a RustFS bucket. The workflow was verified with `apache/kafka:4.0.0` and `confluentinc/kafka-connect-s3` v10.5.25 against `rustfs/rustfs-x86-musl:v2.3.1`. + +Kafka's KIP-405 tiered storage needs a `RemoteLogStorageManager` plugin, and the S3 implementations in the ecosystem are vendor-proprietary. The Connect S3 sink is the open, self-hosted way to move topic data to S3-compatible storage and is the approach documented here. + +You need Docker. This deployment is intended for local integration testing, not production. + +## Architecture + +```mermaid +flowchart LR + Producer["Console producer"] -->|"records"| Broker["Kafka broker :9092"] + Broker -->|"consumer group"| Connect["Connect S3 sink"] + Connect -->|"batched objects"| RustFS["RustFS :9000"] +``` + +The sink task consumes a topic in a dedicated consumer group and writes record batches to the bucket, one object per `flush.size` records per partition. + +## 1. Run the broker + +Start a KRaft broker whose advertised listener is reachable from other containers: + +```bash +docker run -d --name kafka --hostname kafka --network oo-rustfs_default \ + -e CLUSTER_ID=5L6g3nShT-eMCtK--X86sw \ + -e KAFKA_NODE_ID=1 \ + -e KAFKA_PROCESS_ROLES=broker,controller \ + -e KAFKA_LISTENERS=PLAINTEXT://:9092,CONTROLLER://:9093 \ + -e KAFKA_ADVERTISED_LISTENERS=PLAINTEXT://kafka:9092 \ + -e KAFKA_CONTROLLER_LISTENER_NAMES=CONTROLLER \ + -e KAFKA_LISTENER_SECURITY_PROTOCOL_MAP=CONTROLLER:PLAINTEXT,PLAINTEXT:PLAINTEXT \ + -e KAFKA_CONTROLLER_QUORUM_VOTERS=1@kafka:9093 \ + -e KAFKA_OFFSETS_TOPIC_REPLICATION_FACTOR=1 \ + -e KAFKA_TRANSACTION_STATE_LOG_REPLICATION_FACTOR=1 \ + -e KAFKA_TRANSACTION_STATE_LOG_MIN_ISR=1 \ + apache/kafka:4.0.0 + +docker exec kafka /opt/kafka/bin/kafka-topics.sh \ + --bootstrap-server localhost:9092 \ + --create --topic rustfs-topic --partitions 1 --replication-factor 1 + +rc mb rustfs/kafka-demo +``` + +The image defaults to advertising `localhost:9092`, which only works from inside the broker container. The `KAFKA_ADVERTISED_LISTENERS` override is what makes Connect (and any remote client) able to reach the broker. + +## 2. Install the connector + +Download the Confluent Hub archive, which bundles the connector and its dependencies, and unpack it where the worker can see it: + +```bash +curl -Lo kafka-connect-s3.zip "https://hub-downloads.confluent.io/api/plugins/confluentinc/kafka-connect-s3/versions/10.5.25/confluentinc-kafka-connect-s3-10.5.25.zip" +unzip kafka-connect-s3.zip -d /opt/kafka-conn/plugins +``` + +## 3. Configure the worker and the sink + +Create the worker properties. The value converter must be `ByteArrayConverter` so records are written verbatim: + +```ini title="worker.properties" +bootstrap.servers=kafka:9092 +key.converter=org.apache.kafka.connect.storage.StringConverter +value.converter=org.apache.kafka.connect.converters.ByteArrayConverter +offset.storage.file.filename=/tmp/connect.offsets +offset.flush.interval.ms=5000 +plugin.path=/opt/kafka-conn-plugins +``` + +Create the sink connector configuration, replacing all connection placeholders: + +```ini title="rustfs-sink.properties" +name=rustfs-sink +connector.class=io.confluent.connect.s3.S3SinkConnector +tasks.max=1 +topics=rustfs-topic +s3.bucket.name=kafka-demo +s3.region=us-east-1 +store.url=http://:9000 +s3.path.style.access.enabled=true +flush.size=3 +storage.class=io.confluent.connect.s3.storage.S3Storage +format.class=io.confluent.connect.s3.format.bytearray.ByteArrayFormat +consumer.override.auto.offset.reset=earliest +``` + +Kafka 4.0 moved the class to `org.apache.kafka.connect.converters.ByteArrayConverter` — the old `storage` package path no longer resolves. + +## 4. Run the worker + +Run `connect-standalone` in the foreground so its logs go to `docker logs`, with the bucket credentials in the environment: + +```bash +docker run -d --name kafka-connect --hostname kafka-connect \ + --network oo-rustfs_default \ + -v /opt/kafka-conn/plugins:/opt/kafka-conn-plugins:ro \ + -v "$PWD/worker.properties":/etc/kafka/worker.properties:ro \ + -v "$PWD/rustfs-sink.properties":/etc/kafka/sink.properties:ro \ + -e AWS_ACCESS_KEY_ID= \ + -e AWS_SECRET_ACCESS_KEY= \ + apache/kafka:4.0.0 \ + /opt/kafka/bin/connect-standalone.sh /etc/kafka/worker.properties /etc/kafka/sink.properties +``` + +The worker is ready when the sink task claims the partition: + +```text +INFO [rustfs-sink|task-0] Assigned topic partitions: [rustfs-topic-0] +``` + +## 5. Produce records + +Send at least `flush.size` records so the connector completes a batch: + +```bash +docker exec kafka sh -c "printf 'msg-one\nmsg-two\nmsg-three\n' | \ + /opt/kafka/bin/kafka-console-producer.sh --bootstrap-server localhost:9092 --topic rustfs-topic" +``` + +After a few seconds the batch becomes an object in the bucket: + +```bash +rc ls rustfs/kafka-demo/ -r +rc cat rustfs/kafka-demo/topics/rustfs-topic/partition=0/rustfs-topic+0+0000000000.bin +``` + +```text +topics/rustfs-topic/partition=0/rustfs-topic+0+0000000000.bin +topics/rustfs-topic/partition=0/rustfs-topic+0+0000000003.bin +msg-one +msg-two +msg-three +``` + +The object name encodes topic, partition, and starting offset. Each subsequent batch of three records lands in the next object (`+0000000003.bin` and so on). + +![Kafka sink objects stored in the RustFS Console](./images/rustfs-kafka-sink.png) + +## 6. Stop or reset + +To tear down the demo while keeping the bucket objects: + +```bash +docker rm -f kafka-connect kafka +``` + +To delete the stored data: + +```bash +rc rm rustfs/kafka-demo/ --recursive --force +``` + +## Troubleshooting + +### `AdminClient ... Rebootstrapping with Cluster (id: null)` loops forever + +The broker advertises `localhost:9092`, so a remote client receives metadata pointing at itself. Set `KAFKA_ADVERTISED_LISTENERS=PLAINTEXT://kafka:9092` (and matching listener variables) as in step 1. + +### `Invalid schema type for ByteArrayConverter: STRING` + +The `ByteArrayFormat` writer only accepts raw bytes. Either switch the worker's `value.converter` to the ByteArray converter or choose a format class that matches the converter you use. + +### `Class org.apache.kafka.connect.storage.ByteArrayConverter could not be found` + +Kafka 4.0 moved the class to `org.apache.kafka.connect.converters.ByteArrayConverter`. Use the new package path in `worker.properties`. + +### The connector downloads but the plugin is not found + +The plain connector JAR from Maven lacks its dependencies. Use the Confluent Hub archive from step 2, which bundles the complete `lib/` directory, and make sure `plugin.path` points at the directory that contains the connector folder. + +## Next steps + +- Review [S3 compatibility notes](/administration/protocols/s3) before adopting additional connectors. +- Create dedicated production credentials with [Access Key Management](/security-compliance/iam/access-token). +- Follow the [Kafka Connect S3 sink documentation](https://docs.confluent.io/kafka-connect-s3/current/index.html) for Parquet and Avro formats, partitioning by time, and IAM-based credential chains. diff --git a/content/de/developer/integration/big-data/meta.json b/content/de/developer/integration/big-data/meta.json index 29684232..2374fa35 100644 --- a/content/de/developer/integration/big-data/meta.json +++ b/content/de/developer/integration/big-data/meta.json @@ -2,12 +2,15 @@ "title": "Data Analytics", "pages": [ "clickhouse", + "airflow", "duckdb", "doris", + "delta-lake", "flink", "hudi", "iceberg", "influxdb", + "kafka", "lakefs", "milvus", "mlflow", diff --git a/content/de/developer/integration/devops/images/rustfs-opensearch-snapshot.png b/content/de/developer/integration/devops/images/rustfs-opensearch-snapshot.png new file mode 100644 index 00000000..60b3bd3a Binary files /dev/null and b/content/de/developer/integration/devops/images/rustfs-opensearch-snapshot.png differ diff --git a/content/de/developer/integration/devops/index.md b/content/de/developer/integration/devops/index.md index 86b751d4..2b532228 100644 --- a/content/de/developer/integration/devops/index.md +++ b/content/de/developer/integration/devops/index.md @@ -8,6 +8,7 @@ Nutzen Sie **RustFS** als Objektspeicher-Layer für DevOps-Plattformen und Infra ## Plattformen und Tools - [Elasticsearch](./elasticsearch.md) +- [OpenSearch](./opensearch.md) - [Gitea](./gitea.md) - [Jenkins](./jenkins.md) - [Terraform](./terraform.md) diff --git a/content/de/developer/integration/devops/meta.json b/content/de/developer/integration/devops/meta.json index 67c84c6e..55eada6a 100644 --- a/content/de/developer/integration/devops/meta.json +++ b/content/de/developer/integration/devops/meta.json @@ -2,6 +2,7 @@ "title": "DevOps", "pages": [ "elasticsearch", + "opensearch", "gitea", "jenkins", "terraform" diff --git a/content/de/developer/integration/devops/opensearch.md b/content/de/developer/integration/devops/opensearch.md new file mode 100644 index 00000000..9d44e530 --- /dev/null +++ b/content/de/developer/integration/devops/opensearch.md @@ -0,0 +1,172 @@ +--- +title: "OpenSearch" +description: "Snapshot OpenSearch indices to RustFS with the repository-s3 plugin." +--- + +This guide connects [OpenSearch](https://github.com/opensearch-project/OpenSearch) — the open-source search and analytics suite derived from Elasticsearch — to **RustFS** through the `repository-s3` plugin. You will register an S3 snapshot repository backed by a RustFS bucket, take a snapshot of an index, and restore it. The workflow was verified with `opensearchproject/opensearch:3.8.0` and the bundled `repository-s3` plugin against `rustfs/rustfs-x86-musl:v2.3.1`. + +You need Docker, or an OpenSearch node where you can install plugins and edit configuration. This deployment is intended for local integration testing, not production. + +## Architecture + +```mermaid +flowchart LR + Client["REST client"] --> OS["OpenSearch :9200"] + OS -->|"snapshot files"| RustFS["RustFS :9000"] + RustFS -->|"restore"| OS +``` + +The `repository-s3` plugin writes snapshots as shard archives plus metadata blobs in the bucket. Registration is cluster-wide, so every node needs the plugin and the same client configuration. + +## 1. Run OpenSearch + +Start a single node with security disabled and a small heap: + +```bash +docker run -d --name opensearch --network oo-rustfs_default -p 9200:9200 \ + -e discovery.type=single-node \ + -e OPENSEARCH_JAVA_OPTS="-Xms512m -Xmx512m" \ + -e DISABLE_SECURITY_PLUGIN=true \ + opensearchproject/opensearch:3.8.0 +``` + +The node is ready when `curl http://localhost:9200` returns the cluster header (allow one to two minutes). + +## 2. Install the repository-s3 plugin + +The S3 repository plugin is not preloaded. Install it and restart the node: + +```bash +docker exec opensearch bin/opensearch-plugin install --batch repository-s3 +docker restart opensearch +``` + +## 3. Configure the S3 client + +Credentials are secure settings: they belong in the OpenSearch keystore, not in the repository request or `opensearch.yml`. Create the keystore entries, replacing all connection placeholders: + +```bash +docker exec opensearch sh -c \ + "printf '' | bin/opensearch-keystore create 2>/dev/null; \ + printf '' | bin/opensearch-keystore add -f -x s3.client.default.access_key; \ + printf '' | bin/opensearch-keystore add -f -x s3.client.default.secret_key" +``` + +Add the non-secure client settings to `config/opensearch.yml`: + +```yaml title="opensearch.yml" +network.host: 0.0.0.0 +plugins.security.disabled: true +s3.client.default.endpoint: http://:9000 +s3.client.default.protocol: http +s3.client.default.path_style_access: "true" +``` + +Restart the node once more so it reads both the keystore and the new settings: + +```bash +docker restart opensearch +``` + +Create the bucket while the node boots: + +```bash +rc mb rustfs/opensearch-snapshots +``` + +## 4. Register the repository and snapshot + +Create a test index with a document, then register the repository: + +```bash +curl -sX PUT http://localhost:9200/rustfs-demo -H "Content-Type: application/json" \ + -d '{"settings":{"number_of_shards":1}}' + +curl -sX PUT http://localhost:9200/rustfs-demo/_doc/1 -H "Content-Type: application/json" \ + -d '{"product":"rustfs","via":"opensearch-snapshot"}' + +curl -sX PUT "http://localhost:9200/_snapshot/rustfs-repo" -H "Content-Type: application/json" \ + -d '{"type":"s3","settings":{"bucket":"opensearch-snapshots","region":"us-east-1","server_side_encryption_type":"bucket_default"}}' +``` + +The `server_side_encryption_type: bucket_default` setting matters: without it the plugin requests SSE-S3, which a self-hosted RustFS without a server-side encryption master key rejects. + +Take a snapshot and wait for completion: + +```bash +curl -sX PUT "http://localhost:9200/_snapshot/rustfs-repo/snapshot-1?wait_for_completion=true" \ + -H "Content-Type: application/json" -d '{"indices":"rustfs-demo"}' +``` + +```text +{"snapshot":{"snapshot":"snapshot-1","state":"SUCCESS","indices":["rustfs-demo"],...}} +``` + +## 5. Verify objects and restore + +List the bucket: + +```bash +rc ls rustfs/opensearch-snapshots/ -r +``` + +```text +index-0 +index.latest +indices/5x1bwsWaSv2XINIbeoe-RQ/0/__GgxvoCw-TKuMMAYBq5Khag +indices/5x1bwsWaSv2XINIbeoe-RQ/0/snap-kRgFBiuyQIyPo3_-CMp7Hw.dat +meta-kRgFBiuyQIyPo3_-CMp7Hw.dat +snap-kRgFBiuyQIyPo3_-CMp7Hw.dat +``` + +Delete the index and restore it from the snapshot: + +```bash +curl -sX DELETE http://localhost:9200/rustfs-demo +curl -sX POST "http://localhost:9200/_snapshot/rustfs-repo/snapshot-1/_restore?wait_for_completion=true" +curl -s http://localhost:9200/rustfs-demo/_doc/1 +``` + +```text +{"_index":"rustfs-demo","_id":"1","found":true,"_source":{"product":"rustfs","via":"opensearch-snapshot"}} +``` + +![OpenSearch snapshot stored in the RustFS Console](./images/rustfs-opensearch-snapshot.png) + +## 6. Stop or reset + +To tear down the demo while keeping the bucket objects: + +```bash +docker rm -f opensearch +``` + +To delete the stored snapshots: + +```bash +rc rm rustfs/opensearch-snapshots/ --recursive --force +``` + +## Troubleshooting + +### `Setting [access_key] is insecure, but property [allow_insecure_settings] is not set` + +Inline credentials in the repository request are rejected. Store them in the keystore as shown in step 3 — `access_key` and `secret_key` are secure settings in OpenSearch. + +### `SSE-S3 requires RUSTFS_SSE_S3_MASTER_KEY ... (Status Code: 400)` + +The plugin encrypts uploads with SSE-S3 by default. Register the repository with `"server_side_encryption_type": "bucket_default"` so no encryption header is sent, as in step 4. + +### `unknown setting [s3.client.default.access_key]` at startup + +The settings reference the repository-s3 plugin. If the node fails to start with them present, the plugin is not installed in that container — repeat step 2 (a fresh container loses plugins installed with `docker exec`). + +### Repository verification fails with `path is not accessible` + +The node cannot reach the bucket: check that `s3.client.default.endpoint` is reachable from the container, `path_style_access` is `"true"`, and the keystore credentials were loaded (they are read at startup — restart after adding them). + +## Next steps + +- Review [S3 compatibility notes](/administration/protocols/s3) before adopting additional OpenSearch repositories. +- Create dedicated production credentials with [Access Key Management](/security-compliance/iam/access-token). +- Follow the [OpenSearch snapshots documentation](https://docs.opensearch.org/docs/latest/tuning-your-cluster/availability-and-recovery/snapshots/index/) to automate snapshots with Snapshot Management (SM) policies. diff --git a/content/de/developer/integration/index.md b/content/de/developer/integration/index.md index b4d9e222..477aefa7 100644 --- a/content/de/developer/integration/index.md +++ b/content/de/developer/integration/index.md @@ -7,14 +7,14 @@ Use this section to connect **RustFS** to infrastructure and application platfor ## Integration categories -- [Reverse Proxy](./reverse-proxy/index.md) covers Nginx, Traefik, Caddy, and HAProxy. +- [Reverse Proxy](./reverse-proxy/index.md) covers Nginx, Traefik, Caddy, HAProxy, and Envoy. - [Backup](./backup/index.md) covers Kopia, Longhorn, Restic, and Velero. -- [AI](./ai/index.md) covers AI platforms including Ray. -- [Datenanalyse](./big-data/index.md) covers analytics systems including ClickHouse, Doris, Hudi, Iceberg, lakeFS, Milvus, OpenDAL, Vitess, and Zeppelin. +- [AI](./ai/index.md) covers AI platforms including Ray and vLLM. +- [Datenanalyse](./big-data/index.md) covers analytics systems including Airflow, ClickHouse, Delta Lake, Doris, Hudi, Iceberg, Kafka, lakeFS, Milvus, OpenDAL, Vitess, and Zeppelin. - [Cloud Native](./cloud-native/index.md) covers Cortex and Flux. -- [Observability](./observability/index.md) covers telemetry systems including Fluentd, OpenObserve, OpenTelemetry, Thanos, and Tempo. -- [Others](./others/index.md) covers the community-driven capo SDK for Python. +- [Observability](./observability/index.md) covers telemetry systems including Fluentd, GreptimeDB, Loki, OpenObserve, OpenTelemetry, Tempo, Thanos, and VictoriaMetrics. +- [Others](./others/index.md) covers the capo SDK, rclone, JuiceFS, Nextcloud, and tusd. - [Registry](./registry/index.md) covers Harbor. -- [DevOps](./devops/index.md) covers Elasticsearch, Gitea, Jenkins, and Terraform. +- [DevOps](./devops/index.md) covers Elasticsearch, Gitea, Jenkins, OpenSearch, and Terraform. Each guide identifies the RustFS endpoint and addressing requirements to use when configuring the integrating system. \ No newline at end of file diff --git a/content/de/developer/integration/observability/greptimedb.md b/content/de/developer/integration/observability/greptimedb.md new file mode 100644 index 00000000..5b7ee690 --- /dev/null +++ b/content/de/developer/integration/observability/greptimedb.md @@ -0,0 +1,129 @@ +--- +title: "GreptimeDB" +description: "Run GreptimeDB with RustFS as the S3-compatible object storage backend." +--- + +This guide connects [GreptimeDB](https://github.com/GreptimeTeam/greptimedb) — the open-source, cloud-native time-series database — to **RustFS** as its object storage backend. You will start a standalone instance with its `[storage]` section pointed at a RustFS bucket, write time-series rows through the SQL API, and confirm the Parquet files and manifests in the bucket. The workflow was verified with `greptime/greptimedb` (main, commit `179ff8e5`) against `rustfs/rustfs-x86-musl:v2.3.1`. + +You need Docker, or a local GreptimeDB binary. This deployment is intended for local integration testing, not production. + +## Architecture + +```mermaid +flowchart LR + SQL["SQL / Prometheus API"] --> DB["GreptimeDB"] + DB -->|"SST + manifests"| RustFS["RustFS :9000"] +``` + +GreptimeDB keeps its write-ahead log and recent data locally, then persists SSTables (Parquet) and table manifests to object storage. Pointing the storage backend at RustFS makes the bucket the durable home of all table data. + +## 1. Configure the storage backend + +Create the bucket and a config file with an S3 storage section, replacing all connection placeholders: + +```toml title="greptimedb.toml" +[storage] +type = "S3" +bucket = "" +root = "greptimedb" +access_key_id = "" +secret_access_key = "" +endpoint = "http://:9000" +region = "us-east-1" +``` + +GreptimeDB uses path-style requests for custom endpoints by default; virtual-hosted style must be opted into explicitly with `enable_virtual_host_style`, so no extra flag is needed for RustFS. + +## 2. Run GreptimeDB + +Start a standalone instance with the config file: + +```bash +docker run -d --name greptimedb --network oo-rustfs_default -p 4000:4000 -p 4002:4002 \ + -v "$PWD/greptimedb.toml":/etc/greptimedb/greptimedb.toml:ro \ + greptime/greptimedb:latest standalone start \ + --http-addr 0.0.0.0:4000 \ + --mysql-addr 0.0.0.0:4002 \ + --config-file /etc/greptimedb/greptimedb.toml +``` + +Port `4000` serves the HTTP SQL endpoint and `4002` the MySQL protocol. + +## 3. Write and query time series + +Create a table, insert rows, and read them back. The HTTP SQL endpoint takes form-encoded requests: + +```bash +curl -s -X POST "http://localhost:4000/v1/sql" \ + --data-urlencode "sql=CREATE TABLE rustfs_demo (host STRING, cpu DOUBLE, mem DOUBLE, ts TIMESTAMP TIME INDEX)" + +curl -s -X POST "http://localhost:4000/v1/sql" \ + --data-urlencode "sql=INSERT INTO rustfs_demo VALUES (\"node-1\", 0.31, 0.62, 1790681000000), (\"node-1\", 0.35, 0.63, 1790681060000), (\"node-2\", 0.51, 0.71, 1790681000000)" +``` + +```text +{"output":[{"affectedrows":3}],"execution_time_ms":2} +``` + +Query the rows back: + +```bash +curl -s -X POST "http://localhost:4000/v1/sql" \ + --data-urlencode "sql=SELECT * FROM rustfs_demo ORDER BY ts" +``` + +```text +{"output":[{"records":{"rows":[["node-2",0.51,0.71,1790681000000],["node-1",0.35,0.63,1790681060000]],"total_rows":2}}]} +``` + +## 4. Verify objects in RustFS + +List the bucket — after the memtable flushes, the bucket holds Parquet SSTables and JSON manifests: + +```bash +rc ls rustfs// -r +``` + +```text +greptimedb/data/greptime/public/1024/1024_0000000000/manifest/00000000000000000000.json +greptimedb/data/greptime/greptime_private/1025/1025_0000000000/b11e8b25-5763-4f05-bcab-b6ee0a756a69.parquet +greptimedb/data/greptime/greptime_private/1025/1025_0000000000/manifest/00000000000000000001.json +``` + +Each database gets a directory under `data/`, and per-region `manifest/*.json` files describe the SSTables GreptimeDB reads back during queries. + +![GreptimeDB data stored in the RustFS Console](./images/rustfs-greptimedb-data.png) + +## 5. Stop or reset + +To tear down the demo while keeping the bucket objects: + +```bash +docker rm -f greptimedb +``` + +To delete the stored data: + +```bash +rc rm rustfs// --recursive --force +``` + +## Troubleshooting + +### `Form requests must have Content-Type: application/x-www-form-urlencoded` + +The `/v1/sql` HTTP endpoint only accepts form-encoded bodies. Pass SQL with `curl --data-urlencode "sql=..."` (or `application/x-www-form-urlencoded`), not as a JSON body. + +### Bucket stays empty + +GreptimeDB flushes memtables to object storage asynchronously. Run a few more inserts and wait a few seconds, or trigger a manual flush, then list the bucket again. + +### Startup fails with an S3 error + +Confirm `endpoint` includes the scheme, the bucket exists, and `access_key_id`/`secret_access_key` match a RustFS access key. The `root` value is optional but keeps the table tree under a known prefix. + +## Next steps + +- Review [S3 compatibility notes](/administration/protocols/s3) before adopting additional GreptimeDB storage options. +- Create dedicated production credentials with [Access Key Management](/security-compliance/iam/access-token). +- Follow the [GreptimeDB configuration reference](https://docs.greptime.com/operational-guide/configure/configure-datanode/) to tune flush intervals and cache layers for production workloads. diff --git a/content/de/developer/integration/observability/images/rustfs-greptimedb-data.png b/content/de/developer/integration/observability/images/rustfs-greptimedb-data.png new file mode 100644 index 00000000..1874f210 Binary files /dev/null and b/content/de/developer/integration/observability/images/rustfs-greptimedb-data.png differ diff --git a/content/de/developer/integration/observability/images/rustfs-vm-backups.png b/content/de/developer/integration/observability/images/rustfs-vm-backups.png new file mode 100644 index 00000000..a0bd0cee Binary files /dev/null and b/content/de/developer/integration/observability/images/rustfs-vm-backups.png differ diff --git a/content/de/developer/integration/observability/index.md b/content/de/developer/integration/observability/index.md index 99a78611..c9271d57 100644 --- a/content/de/developer/integration/observability/index.md +++ b/content/de/developer/integration/observability/index.md @@ -8,10 +8,12 @@ Nutzen Sie **RustFS** als Objektspeicher-Layer für Observability-Plattformen, d ## Plattformen - [Fluentd](./fluentd.md) +- [GreptimeDB](./greptimedb.md) - [OpenObserve](./openobserve.md) - [OpenTelemetry](./opentelemetry.md) - [Loki](./loki.md) - [Tempo](./tempo.md) - [Thanos](./thanos.md) +- [VictoriaMetrics](./victoriametrics.md) Speichern Sie Telemetriedaten in einem dedizierten Bucket und beschränken Sie die Anmeldeinformationen auf die erforderlichen Bucket-Operationen. diff --git a/content/de/developer/integration/observability/meta.json b/content/de/developer/integration/observability/meta.json index d6ddebd7..97ef4d6e 100644 --- a/content/de/developer/integration/observability/meta.json +++ b/content/de/developer/integration/observability/meta.json @@ -2,10 +2,12 @@ "title": "Observability", "pages": [ "fluentd", + "greptimedb", "loki", "openobserve", "opentelemetry", "tempo", - "thanos" + "thanos", + "victoriametrics" ] } diff --git a/content/de/developer/integration/observability/victoriametrics.md b/content/de/developer/integration/observability/victoriametrics.md new file mode 100644 index 00000000..efc799d9 --- /dev/null +++ b/content/de/developer/integration/observability/victoriametrics.md @@ -0,0 +1,157 @@ +--- +title: "VictoriaMetrics" +description: "Back up VictoriaMetrics snapshots to RustFS with vmbackup." +--- + +This guide connects [VictoriaMetrics](https://github.com/VictoriaMetrics/VictoriaMetrics) — the Prometheus-compatible time-series database — to **RustFS** through `vmbackup` and `vmrestore`. You will run a single-node instance, import metrics, create an instant snapshot, back it up to a RustFS bucket, and restore the data into a fresh directory. The workflow was verified with `victoria-metrics`, `vmbackup`, and `vmrestore` v1.x images against `rustfs/rustfs-x86-musl:v2.3.1`. + +You need Docker. This deployment is intended for local integration testing, not production. + +## Architecture + +```mermaid +flowchart LR + Import["Prometheus import API"] --> VM["VictoriaMetrics :8428"] + VM -->|"instant snapshot"| Backup["vmbackup"] + Backup -->|"S3 API"| RustFS["RustFS :9000"] + RustFS -->|"restore"| Restore["vmrestore"] +``` + +`vmbackup` uploads a consistent point-in-time snapshot of the storage directory to any S3-compatible endpoint. `vmrestore` reverses the process, producing a data directory a VictoriaMetrics instance can open directly. + +## 1. Run VictoriaMetrics + +Create the bucket and start a single-node instance: + +```bash +rc mb rustfs/vm-backups + +docker run -d --name vm --network oo-rustfs_default -p 8428:8428 \ + -v vm-data:/storage \ + victoriametrics/victoria-metrics:latest \ + -storageDataPath=/storage -retentionPeriod=100y +``` + +## 2. Import metrics + +Write a couple of samples through the Prometheus import API: + +```bash +echo "vm_demo_metric 123" | curl -s --data-binary @- http://localhost:8428/api/v1/import/prometheus +echo "vm_demo_metric 456" | curl -s --data-binary @- http://localhost:8428/api/v1/import/prometheus +``` + +The endpoint answers `204 No Content`. Confirm the data is queryable: + +```bash +curl -s "http://localhost:8428/api/v1/export?match[]=vm_demo_metric" +``` + +```text +{"metric":{"__name__":"vm_demo_metric"},"values":[123,456],"timestamps":[1790680894604,1790680894619]} +``` + +## 3. Create a snapshot + +Ask VictoriaMetrics for a consistent snapshot: + +```bash +curl -s http://localhost:8428/snapshot/create +``` + +```text +{"status":"ok","snapshot":"20260929112134-18D9C6C76704F913"} +``` + +## 4. Back the snapshot up to RustFS + +Run `vmbackup` against the same storage volume, replacing the credential placeholders. The snapshot name comes from step 3: + +```bash +docker run --rm --network oo-rustfs_default \ + -e AWS_ACCESS_KEY_ID= \ + -e AWS_SECRET_ACCESS_KEY= \ + --volumes-from vm \ + victoriametrics/vmbackup:latest \ + -storageDataPath=/storage \ + -snapshotName=20260929112134-18D9C6C76704F913 \ + -dst=s3://vm-backups/demo \ + -customS3Endpoint=http://:9000 +``` + +```text +backup ... to S3{bucket: "vm-backups", dir: "demo/"} is complete; uploaded 760 bytes +``` + +`-customS3Endpoint` redirects the AWS SDK to RustFS; custom endpoints are addressed with path-style requests automatically. Set `AWS_EC2_METADATA_DISABLED=true` on hosts without an EC2 metadata service to skip credential lookup delays. + +## 5. Verify and restore + +List the bucket prefix: + +```bash +rc ls rustfs/vm-backups/demo/ +``` + +```text +backup_complete.ignore +backup_metadata.ignore +data/ +metadata/ +``` + +`backup_complete.ignore` marks a complete backup. Restore it into a fresh directory: + +```bash +docker run --rm --network oo-rustfs_default \ + -e AWS_ACCESS_KEY_ID= \ + -e AWS_SECRET_ACCESS_KEY= \ + -v /opt/vm-restore:/restore \ + victoriametrics/vmrestore:latest \ + -src=s3://vm-backups/demo \ + -storageDataPath=/restore \ + -customS3Endpoint=http://:9000 +``` + +```text +restored 760 bytes from backup in 0.055 seconds +``` + +The restored directory contains `data/`, `metadata/`, and a lock file — exactly what a VictoriaMetrics instance expects at `-storageDataPath`. + +![VictoriaMetrics backup stored in the RustFS Console](./images/rustfs-vm-backups.png) + +## 6. Stop or reset + +To tear down the demo while keeping the bucket objects: + +```bash +docker rm -f vm +docker volume rm vm-data +``` + +To delete the stored backups: + +```bash +rc rm rustfs/vm-backups/ --recursive --force +``` + +## Troubleshooting + +### `vmbackup` hangs at startup or fails to find credentials + +The AWS SDK probes the EC2 metadata service when environment credentials are absent. On machines without IMDS, export `AWS_EC2_METADATA_DISABLED=true` next to the key variables. + +### Backup parts re-upload on every run + +`vmbackup` performs incremental backups by comparing local and remote file hashes. Restoring to a fresh directory and running `vmbackup` from there re-uploads everything; keep the original data directory for incremental runs. + +### Query returns nothing right after import + +Imports are accepted asynchronously and the instant query endpoint can lag on a busy single node. Verify with `/api/v1/export` (or wait a few seconds) before creating the snapshot. + +## Next steps + +- Review [S3 compatibility notes](/administration/protocols/s3) before adopting additional VictoriaMetrics components. +- Create dedicated production credentials with [Access Key Management](/security-compliance/iam/access-token). +- Follow the [vmbackup documentation](https://docs.victoriametrics.com/vmbackup/) to schedule backups and prune old snapshots. diff --git a/content/de/developer/integration/others/images/rustfs-juicefs-chunks.png b/content/de/developer/integration/others/images/rustfs-juicefs-chunks.png new file mode 100644 index 00000000..737490b6 Binary files /dev/null and b/content/de/developer/integration/others/images/rustfs-juicefs-chunks.png differ diff --git a/content/de/developer/integration/others/images/rustfs-nextcloud-file.png b/content/de/developer/integration/others/images/rustfs-nextcloud-file.png new file mode 100644 index 00000000..506e0d79 Binary files /dev/null and b/content/de/developer/integration/others/images/rustfs-nextcloud-file.png differ diff --git a/content/de/developer/integration/others/images/rustfs-rclone-sync.png b/content/de/developer/integration/others/images/rustfs-rclone-sync.png new file mode 100644 index 00000000..dcf13712 Binary files /dev/null and b/content/de/developer/integration/others/images/rustfs-rclone-sync.png differ diff --git a/content/de/developer/integration/others/images/rustfs-tus-uploads.png b/content/de/developer/integration/others/images/rustfs-tus-uploads.png new file mode 100644 index 00000000..e82a55e1 Binary files /dev/null and b/content/de/developer/integration/others/images/rustfs-tus-uploads.png differ diff --git a/content/de/developer/integration/others/index.md b/content/de/developer/integration/others/index.md index 9b1e3603..c4d3c2cf 100644 --- a/content/de/developer/integration/others/index.md +++ b/content/de/developer/integration/others/index.md @@ -8,3 +8,7 @@ Integration guides that do not fit the other categories. ## Guides - [capo (Python)](./capo.md) — connect the community-driven capo SDK to RustFS with synchronous or asynchronous clients. +rclone](./rclone.md) — sync, mount, and serve RustFS buckets from the command line. +- [tusd](./tusd.md) — receive resumable uploads into a RustFS bucket over the tus protocol. +- [JuiceFS](./juicefs.md) — mount a POSIX filesystem backed by a RustFS bucket. +- [Nextcloud](./nextcloud.md) — use RustFS as S3 external storage for Nextcloud files. diff --git a/content/de/developer/integration/others/juicefs.md b/content/de/developer/integration/others/juicefs.md new file mode 100644 index 00000000..428b08ce --- /dev/null +++ b/content/de/developer/integration/others/juicefs.md @@ -0,0 +1,132 @@ +--- +title: "JuiceFS" +description: "Build a POSIX filesystem on RustFS with JuiceFS S3 object storage." +--- + +This guide connects [JuiceFS](https://github.com/juicedata/juicefs) — the cloud-native distributed POSIX filesystem — to **RustFS** as its object storage backend. You will format a volume whose data chunks live in a RustFS bucket, mount it locally, and read and write files through the mount. The workflow was verified with `juicefs v1.3.1` (community edition, SQLite metadata engine) against `rustfs/rustfs-x86-musl:v2.3.1`. + +You need the JuiceFS binary, a metadata engine, and FUSE (`fuse3` on Linux). This deployment is intended for local integration testing, not production. + +## Architecture + +```mermaid +flowchart LR + Mount["/mnt/jfs"] -->|"POSIX"| JuiceFS["JuiceFS client"] + JuiceFS -->|"metadata"| Meta["SQLite / Redis"] + JuiceFS -->|"data chunks"| RustFS["RustFS :9000"] +``` + +JuiceFS splits every file into chunks and stores them as objects under `chunks/` in the bucket, while the metadata engine tracks names, inodes, and layout. The filesystem behaves like a local disk but holds no data locally. + +## 1. Format the volume + +Create the bucket and format a JuiceFS volume backed by RustFS, replacing all connection placeholders. The bucket URL carries the endpoint, which selects path-style addressing: + +```bash +rc mb rustfs/jfs-demo + +juicefs format \ + --storage s3 \ + --bucket http://:9000/jfs-demo \ + --access-key \ + --secret-key \ + sqlite3:///opt/juicefs/jfs.db \ + rustfs-jfs +``` + +```text +Data use s3://:9000/jfs-demo/rustfs-jfs/ + OK, rustfs-jfs is ready +``` + +`sqlite3:///opt/juicefs/jfs.db` is the metadata engine for this test. In production, use Redis, MySQL, or PostgreSQL instead so multiple clients can mount the same volume. + +## 2. Mount the volume + +Mount the filesystem with the same metadata URL: + +```bash +mkdir -p /mnt/jfs +juicefs mount -d sqlite3:///opt/juicefs/jfs.db /mnt/jfs +``` + +```text +OK, rustfs-jfs is ready at /mnt/jfs +``` + +The `-d` flag runs the mount in the background. The volume is now a POSIX filesystem. + +## 3. Read and write files + +Use the mount like any other directory: + +```bash +echo "hello rustfs jfs" > /mnt/jfs/hello.txt +dd if=/dev/urandom of=/mnt/jfs/blob.bin bs=1M count=3 +mkdir -p /mnt/jfs/dir1 && echo nested > /mnt/jfs/dir1/nested.txt +cat /mnt/jfs/hello.txt +``` + +```text +hello rustfs jfs +``` + +Inspect the volume with `juicefs info`: + +```text +/mnt/jfs : + inode: 1 + files: 2 + dirs: 2 + length: 3.00 MiB +``` + +## 4. Verify chunks in RustFS + +List the bucket prefixes: + +```bash +rc ls rustfs/jfs-demo/ -r +``` + +Every file was split into content-addressed chunk objects: + +```text +rustfs-jfs/chunks/0/0/1_0_17 +rustfs-jfs/chunks/0/0/3_0_3145728 +rustfs-jfs/chunks/0/0/4_0_7 +``` + +The chunk name encodes the inode, chunk index, and size — for example `3_0_3145728` is the 3 MiB file written in step 3. + +![JuiceFS data chunks stored in the RustFS Console](./images/rustfs-juicefs-chunks.png) + +## 5. Stop or reset + +Unmount the volume, then optionally wipe the volume metadata and bucket data: + +```bash +juicefs umount /mnt/jfs +juicefs destroy --force sqlite3:///opt/juicefs/jfs.db rustfs-jfs +rc rm rustfs/jfs-demo/ --recursive --force +``` + +## Troubleshooting + +### `unknown option: --daemon` + +The background flag is a single dash: `juicefs mount -d`. Running without it keeps the mount in the foreground (useful for debugging). + +### `fusermount3: mount failed: Permission denied` + +Mounting requires the FUSE device. Inside a container, add `--device /dev/fuse --cap-add SYS_ADMIN` (or `--privileged`); on a host, install `fuse3` and confirm `/dev/fuse` exists. + +### Mount hangs or fails to reach storage + +The client must reach both the metadata engine and the bucket endpoint. Because the bucket URL embeds the endpoint, verify it from the mounting host with `curl` before formatting. + +## Next steps + +- Review [S3 compatibility notes](/administration/protocols/s3) before adopting additional JuiceFS storage backends. +- Create dedicated production credentials with [Access Key Management](/security-compliance/iam/access-token). +- Follow the [JuiceFS documentation](https://juicefs.com/docs/community/quick_start_guide/) to switch the metadata engine to Redis and mount the volume from multiple clients. diff --git a/content/de/developer/integration/others/meta.json b/content/de/developer/integration/others/meta.json index f80dfcb6..aeb90a3f 100644 --- a/content/de/developer/integration/others/meta.json +++ b/content/de/developer/integration/others/meta.json @@ -1,6 +1,10 @@ { "title": "Sonstige", "pages": [ - "capo" + "capo", + "juicefs", + "nextcloud", + "rclone", + "tusd" ] } diff --git a/content/de/developer/integration/others/nextcloud.md b/content/de/developer/integration/others/nextcloud.md new file mode 100644 index 00000000..f506d75c --- /dev/null +++ b/content/de/developer/integration/others/nextcloud.md @@ -0,0 +1,145 @@ +--- +title: "Nextcloud" +description: "Use RustFS as S3 external storage for Nextcloud files." +--- + +This guide connects [Nextcloud](https://github.com/nextcloud/server) — the self-hosted content collaboration platform — to **RustFS** through its External Storage app with the S3 backend. You will enable `files_external`, mount a RustFS bucket into every user's files view, and upload a file through WebDAV that lands directly in the bucket. The workflow was verified with `nextcloud:32.0.15` (SQLite, single container) against `rustfs/rustfs-x86-musl:v2.3.1`. + +You need Docker, or an existing Nextcloud instance with `occ` access. This deployment is intended for local integration testing, not production. + +## Architecture + +```mermaid +flowchart LR + User["Browser / WebDAV"] --> Nextcloud["Nextcloud"] + Nextcloud -->|"files_external (S3)"| RustFS["RustFS :9000"] +``` + +Nextcloud proxies file operations on the mount point to the S3 backend. Objects are stored under their mount-relative paths, so the bucket mirrors the names users see. + +## 1. Install Nextcloud + +Run Nextcloud with an admin account, replacing all connection placeholders. SQLite keeps the test self-contained; use MariaDB or PostgreSQL in production: + +```bash +docker run -d --name nextcloud --network oo-rustfs_default -p 8080:80 \ + -e NEXTCLOUD_ADMIN_USER= \ + -e NEXTCLOUD_ADMIN_PASSWORD= \ + nextcloud:32.0.15 +``` + +If the web UI still shows the installer after startup, finish it manually: + +```bash +docker exec -u www-data nextcloud php occ maintenance:install \ + --admin-user --admin-password +``` + +## 2. Enable the External Storage app + +The `files_external` app ships with Nextcloud but starts disabled, and its `occ` commands only exist once the app is enabled: + +```bash +docker exec -u www-data nextcloud php occ app:enable files_external +``` + +```text +files_external 1.24.1 enabled +``` + +## 3. Mount the RustFS bucket + +Create an external storage of backend type `amazons3` with the `amazons3::accesskey` authentication backend. Replace all connection placeholders: + +```bash +docker exec -u www-data nextcloud php occ files_external:create \ + /rustfs amazons3 amazons3::accesskey \ + --user \ + --config bucket= \ + --config hostname= \ + --config port=9000 \ + --config use_ssl=false \ + --config use_path_style=true \ + --config key= \ + --config secret= +``` + +```text +Storage created with id 1 +``` + +The mount point `/rustfs` appears in the files view of the given user. `use_path_style=true` is required for a non-AWS endpoint. Check the connection before using it: + +```bash +docker exec -u www-data nextcloud php occ files_external:verify 1 +``` + +```text + - status: ok + - code: 0 +``` + +## 4. Upload a file and verify + +Upload through the WebDAV endpoint, which writes through the external storage: + +```bash +echo "nextcloud writes to rustfs" > /tmp/nc-demo.txt + +curl -u : \ + -T /tmp/nc-demo.txt \ + http://localhost:8080/remote.php/dav/files//rustfs/nc-demo.txt \ + -o /dev/null -w "%{http_code}\n" +``` + +```text +201 +``` + +Read it back through the same path, then confirm the object in RustFS: + +```bash +rc ls rustfs// -r +``` + +```text +[2026-09-29 13:42:48] 27 B nc-demo.txt +``` + +The object key equals the path inside the mount, so files uploaded through Nextcloud can also be read directly with any S3 client. + +![Nextcloud file stored in the RustFS Console](./images/rustfs-nextcloud-file.png) + +## 5. Stop or reset + +To remove the mount without touching the bucket: + +```bash +docker exec -u www-data nextcloud php occ files_external:delete 1 +``` + +To delete the bucket contents: + +```bash +rc rm rustfs// --recursive --force +``` + +## Troubleshooting + +### `There are no commands defined in the "files_external" namespace` + +The app is not enabled yet. Run `occ app:enable files_external` first; the `occ files_external:*` commands only register afterwards. + +### `Not enough arguments (missing: "authentication_backend")` + +`files_external:create` takes the storage backend and the authentication backend as two separate arguments: `amazons3 amazons3::accesskey`. The backend identifiers are listed by `occ files_external:backends`. + +### Mount shows but is empty, or uploads fail + +Confirm `hostname` is reachable from the Nextcloud container (use the container network name, not `localhost`), `use_path_style` is `true`, and the bucket exists. `occ files_external:verify ` reports the exact connection error. + +## Next steps + +- Review [S3 compatibility notes](/administration/protocols/s3) before adopting additional external storage backends. +- Create dedicated production credentials with [Access Key Management](/security-compliance/iam/access-token). +- Follow the [Nextcloud external storage documentation](https://docs.nextcloud.com/server/latest/admin_manual/configuration_files/external_storage_configuration_gui.html) to share the mount with groups and enable versioning. diff --git a/content/de/developer/integration/others/rclone.md b/content/de/developer/integration/others/rclone.md new file mode 100644 index 00000000..34f9fe4a --- /dev/null +++ b/content/de/developer/integration/others/rclone.md @@ -0,0 +1,144 @@ +--- +title: "rclone" +description: "Sync, mount, and serve RustFS buckets with rclone over its S3-compatible API." +--- + +This guide connects [rclone](https://github.com/rclone/rclone) — the command-line tool for syncing files to and from cloud storage — to **RustFS** through its S3 backend. You will configure an S3 remote for RustFS, copy and sync files, read objects back, publish a bucket over HTTP with `rclone serve`, and mount the bucket as a local filesystem with `rclone mount`. The workflow was verified with `rclone v1.75.1` against `rustfs/rustfs-x86-musl:v2.3.1`. + +You need Docker, or a local rclone binary. This deployment is intended for local integration testing, not production. + +## Architecture + +```mermaid +flowchart LR + Files["Local files"] -->|"copy / sync"| Remote["rclone S3 remote"] + Remote -->|"S3 API"| RustFS["RustFS :9000"] + RustFS -->|"mount / serve"| Client["FUSE mount / HTTP clients"] +``` + +One remote definition drives every rclone command: data transfer, mounting, and serving all use the same S3 connection. + +## 1. Configure the remote + +Create an rclone config file with an S3 remote for RustFS, replacing all connection placeholders. The `Other` provider disables AWS-specific behavior, and path-style addressing is used automatically for custom endpoints: + +```ini title="rclone.conf" +[rustfs] +type = s3 +provider = Other +access_key_id = +secret_access_key = +endpoint = http://:9000 +region = us-east-1 +``` + +## 2. Copy and read objects + +Create the bucket and upload a directory with `rclone copy`: + +```bash +rc mb rustfs/rclone-demo +rclone copy /data rustfs:rclone-demo/seed +``` + +List and read back: + +```bash +rclone ls rustfs:rclone-demo/seed +rclone cat rustfs:rclone-demo/seed/hello.txt +``` + +```text + 3145728 blob.bin + 18 hello.txt +hello from rclone +``` + +`rclone lsd rustfs:` lists every bucket on the endpoint. + +## 3. Sync a directory + +`rclone sync` makes the destination identical to the source, including deletions. Remove a local file and sync: + +```bash +rm /data/hello.txt +rclone sync /data rustfs:rclone-demo/seed +rclone lsf rustfs:rclone-demo/seed +``` + +```text +blob.bin +``` + +`hello.txt` disappears from the bucket. Add `--dry-run` first to preview the changes without touching the bucket. + +## 4. Serve a bucket over HTTP + +Publish the bucket contents as an HTTP file server: + +```bash +rclone serve http --addr 0.0.0.0:8080 rustfs:rclone-demo/seed +``` + +Any HTTP client can now download objects: + +```bash +curl -s http://localhost:8080/blob.bin -o /dev/null -w "%{http_code} %{size_download} bytes\n" +``` + +```text +200 3145728 bytes +``` + +`rclone serve` also supports WebDAV, SFTP, and S3 endpoints over the same remote. + +## 5. Mount the bucket as a filesystem + +With FUSE available, mount the bucket locally and use it like a directory: + +```bash +rclone mount rustfs:rclone-demo /mnt/rclone --daemon +ls /mnt/rclone/seed +echo test > /mnt/rclone/write-test.txt +cat /mnt/rclone/write-test.txt +``` + +Files written through the mount appear in RustFS as regular objects: + +```bash +rc ls rustfs/rclone-demo/ -r +``` + +```text +[2026-09-29 11:12:41] 5 B write-test.txt +``` + +Unmount with `fusermount -u /mnt/rclone` when finished. + +## 6. Stop or reset + +rclone holds no server-side state. To delete the demo data: + +```bash +rclone purge rustfs:rclone-demo +``` + +## Troubleshooting + +### `Access Denied` or empty listings + +Confirm `endpoint` includes the scheme and that the key pair matches a RustFS access key. The `region` value is required by the S3 signer even though RustFS ignores it; keep `us-east-1`. + +### Mount fails with `fusermount3: mount failed: Permission denied` + +Mounting needs the FUSE device and elevated privileges. Inside a container, run with `--device /dev/fuse --cap-add SYS_ADMIN` — and use `--privileged` if the mount helper still fails. On a host, verify that `fuse3` is installed and `/dev/fuse` exists. + +### Sync deleted nothing on the destination + +`rclone copy` never deletes. Only `rclone sync` (or `rclone delete`) removes destination objects, and `--dry-run` is the safe way to preview either. + +## Next steps + +- Review [S3 compatibility notes](/administration/protocols/s3) before adopting additional rclone backends. +- Create dedicated production credentials with [Access Key Management](/security-compliance/iam/access-token). +- Follow the [rclone S3 documentation](https://rclone.org/s3/) for flags such as `--transfers`, bandwidth limits, and crypt overlays. diff --git a/content/de/developer/integration/others/tusd.md b/content/de/developer/integration/others/tusd.md new file mode 100644 index 00000000..99f3c588 --- /dev/null +++ b/content/de/developer/integration/others/tusd.md @@ -0,0 +1,170 @@ +--- +title: "tusd" +description: "Receive resumable uploads into RustFS with the tusd server's S3 backend." +--- + +This guide connects [tusd](https://github.com/tus/tusd) — the official reference implementation of the tus resumable-upload protocol — to **RustFS** as its S3 storage backend. You will run tusd against a RustFS bucket, create an upload with the tus protocol, send the file in two chunks with an interruption in between, resume from the reported offset, and verify the assembled object in the bucket. The workflow was verified with `tusproject/tusd:v2.10.1` against `rustfs/rustfs-x86-musl:v2.3.1`. + +You need Docker, or a local tusd binary. This deployment is intended for local integration testing, not production. + +## Architecture + +```mermaid +flowchart LR + Client["tus client"] -->|"POST / PATCH / HEAD"| tusd["tusd :8080"] + tusd -->|"multipart upload"| RustFS["RustFS :9000"] +``` + +tusd stores each in-progress upload as S3 multipart parts in the bucket. A client that loses its connection asks the server for the last committed offset with `HEAD` and continues from there — the data already received is never sent twice. + +## 1. Run tusd + +Create the bucket and start tusd with the S3 backend, replacing all connection placeholders. The AWS region must be set even though RustFS ignores it: + +```bash +rc mb rustfs/tus-uploads + +docker run -d --name tusd --network oo-rustfs_default -p 8080:8080 \ + -e AWS_ACCESS_KEY_ID= \ + -e AWS_SECRET_ACCESS_KEY= \ + -e AWS_REGION=us-east-1 \ + tusproject/tusd:latest \ + -s3-bucket tus-uploads \ + -s3-endpoint http://:9000 +``` + +Check that the server is healthy: + +```bash +curl -s -o /dev/null -w "%{http_code}\n" http://localhost:8080/health +``` + +```text +200 +``` + +## 2. Create the upload + +Create a 6 MiB upload and read the `Location` header: + +```bash +curl -s -D - -o /dev/null -X POST http://localhost:8080/files/ \ + -H "Upload-Length: 6291456" -H "Tus-Resumable: 1.0.0" \ + | grep -i "^Location:" +``` + +```text +Location: http://localhost:8080/files/b7338250daa9...+NGVmNDRhZjEt... +``` + +The upload URL contains the file ID and a message-authentication tag. Strip the scheme and host before re-sending it (the server echoes whatever `Host` it received, which may not be reachable from your next client). + +## 3. Upload in chunks with an interruption + +Send the first 2.5 MB, then stop — this is the point where a mobile client would lose its connection: + +```bash +head -c 2500000 demo.bin > part1.bin + +curl -s -o /dev/null -w "%{http_code}\n" -X PATCH "http://localhost:8080${LOC}" \ + -H "Upload-Offset: 0" -H "Tus-Resumable: 1.0.0" \ + -H "Content-Type: application/offset+octet-stream" \ + --data-binary @part1.bin +``` + +```text +204 +``` + +Ask the server how much it actually has — this is the resumable-upload core: + +```bash +curl -s -X HEAD "http://localhost:8080${LOC}" \ + -H "Tus-Resumable: 1.0.0" -D - -o /dev/null | grep -i upload-offset +``` + +```text +Upload-Offset: 2500000 +``` + +## 4. Resume and finish + +Continue from offset 2500000 with the remaining bytes: + +```bash +tail -c 3791456 demo.bin > part2.bin + +curl -s -o /dev/null -w "%{http_code}\n" -X PATCH "http://localhost:8080${LOC}" \ + -H "Upload-Offset: 2500000" -H "Tus-Resumable: 1.0.0" \ + -H "Content-Type: application/offset+octet-stream" \ + --data-binary @part2.bin +``` + +```text +204 +``` + +Download the finished upload through tusd and compare checksums with the source: + +```bash +curl -s -o download.bin "http://localhost:8080${LOC}" +sha1sum demo.bin download.bin +``` + +```text +d9016032ced6c7515b67a0c556e006c4b25a5858 demo.bin +d9016032ced6c7515b67a0c556e006c4b25a5858 download.bin +``` + +## 5. Verify objects in RustFS + +List the bucket: + +```bash +rc ls rustfs/tus-uploads/ -r +``` + +The bucket holds the assembled object plus one `.info` metadata file per upload — both live and finished: + +```text +b7338250daa9a1a79c1343502b57b28f 6 MiB +b7338250daa9a1a79c1343502b57b28f.info 378 B +``` + +The object key is the upload ID, and the object body is the uploaded file byte-for-byte — so any S3 client can read completed uploads directly from the bucket. + +![tus uploads stored in the RustFS Console](./images/rustfs-tus-uploads.png) + +## 6. Stop or reset + +To tear down the server while keeping the bucket objects: + +```bash +docker rm -f tusd +``` + +To delete the stored uploads: + +```bash +rc rm rustfs/tus-uploads/ --recursive --force +``` + +## Troubleshooting + +### `CreateMultipartUpload ... A region must be set when sending requests to S3` + +tusd builds its S3 client from the AWS environment, and the region is mandatory for endpoint resolution. Export `AWS_REGION=us-east-1` next to the credentials, as in step 1. + +### `PATCH` returns `404` or connects to the wrong host + +The `Location` URL echoes the `Host` header of the creation request. When your client and the server use different hostnames (container name versus published port), strip the scheme and host from the URL and send the path to the address the client can reach. + +### Upload disappears after server restart + +The S3 backend keeps `.info` files in the bucket, so uploads survive restarts. If you run tusd against an empty bucket that another process prunes, the metadata is lost — protect the `tus-uploads` prefix from cleanup jobs. + +## Next steps + +- Review [S3 compatibility notes](/administration/protocols/s3) before adopting additional tusd backends. +- Create dedicated production credentials with [Access Key Management](/security-compliance/iam/access-token). +- Follow the [tus protocol documentation](https://tus.io/protocols/resumable-upload) for creation-with-upload, termination, and checksum extensions that tusd supports on top of the core protocol. diff --git a/content/de/developer/integration/reverse-proxy/envoy.md b/content/de/developer/integration/reverse-proxy/envoy.md new file mode 100644 index 00000000..dc635a48 --- /dev/null +++ b/content/de/developer/integration/reverse-proxy/envoy.md @@ -0,0 +1,192 @@ +--- +title: "Envoy" +description: "Deploy RustFS behind Envoy with TLS-terminated routes for the S3 API and Console." +--- + +Use **Envoy** to terminate TLS and route separate hostnames to the RustFS S3 API and Console. This deployment runs Envoy and a single-node RustFS instance on one Docker network. You need Docker Engine, two DNS records, and a TLS certificate that covers both hostnames (the guide uses a self-signed certificate for testing). + +This guide uses these example hostnames: + +- `s3.example.com` for the S3 API +- `console.example.com` for the Console + +Replace them with hostnames that resolve to the Docker host. + +:::warning[Serve S3 from the root path] + +Do not publish the S3 API under a path such as `/s3/`. AWS Signature Version 4 includes the request path and host, so rewriting either value can invalidate signed requests. Envoy forwards the incoming `Host` header unchanged, which keeps signatures valid. + +::: + +## 1. Create the deployment directories + +Create directories for the Envoy configuration and TLS certificate: + +```bash +mkdir -p rustfs-envoy/certs +cd rustfs-envoy +``` + +For local testing, generate a self-signed certificate covering both hostnames: + +```bash +openssl req -x509 -newkey rsa:2048 -nodes \ + -keyout certs/privkey.pem -out certs/fullchain.pem -days 30 \ + -subj "/CN=*.example.com" \ + -addext "subjectAltName=DNS:s3.example.com,DNS:console.example.com" +chmod 644 certs/privkey.pem +``` + +Make the certificate readable by the non-root user the Envoy image runs as — a `600` private key produces a misleading `Failed to load incomplete private key` error. + +## 2. Configure Envoy + +Create the configuration with an HTTPS listener and two virtual hosts. Route timeouts are disabled (`timeout: 0s`) so long streaming S3 uploads are not cut off: + +```yaml title="envoy.yaml" +static_resources: + listeners: + - name: https + address: {socket_address: {address: 0.0.0.0, port_value: 8443}} + filter_chains: + - transport_socket: + name: envoy.transport_sockets.tls + typed_config: + "@type": type.googleapis.com/envoy.extensions.transport_sockets.tls.v3.DownstreamTlsContext + common_tls_context: + tls_certificates: + - certificate_chain: {filename: /certs/fullchain.pem} + private_key: {filename: /certs/privkey.pem} + filters: + - name: envoy.filters.network.http_connection_manager + typed_config: + "@type": type.googleapis.com/envoy.extensions.filters.network.http_connection_manager.v3.HttpConnectionManager + stat_prefix: rustfs_https + route_config: + virtual_hosts: + - name: s3 + domains: ["s3.example.com", "s3.example.com:*"] + routes: + - match: {prefix: "/"} + route: {cluster: rustfs_s3, timeout: 0s} + - name: console + domains: ["console.example.com", "console.example.com:*"] + routes: + - match: {prefix: "/"} + route: {cluster: rustfs_console, timeout: 0s} + http_filters: + - name: envoy.filters.http.router + typed_config: + "@type": type.googleapis.com/envoy.extensions.filters.http.router.v3.Router + clusters: + - name: rustfs_s3 + connect_timeout: 5s + type: STRICT_DNS + lb_policy: ROUND_ROBIN + load_assignment: + cluster_name: rustfs_s3 + endpoints: + - lb_endpoints: + - endpoint: {address: {socket_address: {address: rustfs, port_value: 9000}}} + - name: rustfs_console + connect_timeout: 5s + type: STRICT_DNS + lb_policy: ROUND_ROBIN + load_assignment: + cluster_name: rustfs_console + endpoints: + - lb_endpoints: + - endpoint: {address: {socket_address: {address: rustfs, port_value: 9001}}} +``` + +The `host:*` domain entries matter: clients send `Host: s3.example.com:8443` on non-standard ports, and Envoy matches the authority including the port. + +## 3. Start Envoy + +Run Envoy on the same Docker network as RustFS, publishing only the proxy port: + +```bash +docker run -d --name envoy --network oo-rustfs_default -p 8443:8443 \ + -v "$PWD/envoy.yaml":/envoy.yaml:ro \ + -v "$PWD/certs":/certs:ro \ + envoyproxy/envoy:v1.34-latest -c /envoy.yaml +``` + +## 4. Verify both endpoints + +Point the example hostnames at the proxy with `curl --resolve` (in production, DNS does this): + +```bash +curl -sk --resolve s3.example.com:8443:127.0.0.1 \ + https://s3.example.com:8443/health/ready -o /dev/null -w "s3 api: %{http_code}\n" + +curl -sk --resolve console.example.com:8443:127.0.0.1 \ + https://console.example.com:8443/rustfs/console/ -o /dev/null -w "console: %{http_code}\n" +``` + +```text +s3 api: 200 +console: 200 +``` + +`-k` skips certificate validation because the certificate is self-signed; with a trusted certificate, drop it. + +## 5. Send signed S3 requests through Envoy + +Point any S3 client at the proxy as if it were RustFS. Configure the client with `https://s3.example.com:8443` as the endpoint and path-style addressing; when the proxy certificate is trusted, signed AWS Signature Version 4 requests pass through unchanged. For a quick test over plain HTTP, add an HTTP listener on port 8080 with the same virtual-host routing as the HTTPS listener, then use the endpoint `http://s3.example.com:8080`: + +```bash +rc alias set rustfs-envoy http://s3.example.com:8080 +rc ls rustfs-envoy/rclone-demo/ +``` + +```text +[ ] 0B seed/ +[2026-09-29 11:12:41] 5 B write-test.txt +``` + +The signatures validate because Envoy forwards the original `Host` header to RustFS. + +## Multi-node backends + +For a distributed RustFS deployment, add every node to the S3 cluster: + +```yaml title="envoy.yaml" + - name: rustfs_s3 + connect_timeout: 5s + type: STRICT_DNS + lb_policy: ROUND_ROBIN + load_assignment: + cluster_name: rustfs_s3 + endpoints: + - lb_endpoints: + - endpoint: {address: {socket_address: {address: node1, port_value: 9000}}} + - lb_endpoints: + - endpoint: {address: {socket_address: {address: node2, port_value: 9000}}} + - lb_endpoints: + - endpoint: {address: {socket_address: {address: node3, port_value: 9000}}} + - lb_endpoints: + - endpoint: {address: {socket_address: {address: node4, port_value: 9000}}} +``` + +The Console cluster follows the same pattern on port `9001`. + +## Troubleshooting + +### `Failed to load incomplete private key from path` + +The Envoy container runs as a non-root user and cannot read a `600` root-owned key. `chmod 644` the key files (or chown them to the container user, UID `1001` in the official image). + +### Routes return `404` with the correct hostnames + +Envoy matches the authority including the port. Add the `host:*` variants to each virtual host's `domains` list, as in the configuration above. + +### `Access Denied` from RustFS on proxied requests + +Confirm the proxy is not rewriting the path or the `Host` header. Signed requests must reach RustFS with the host the client signed for. + +## Next steps + +- [Configure an S3 client](/developer/examples/aws-cli) +- [Enable virtual-hosted-style bucket URLs](/integration/virtual) +- [Review health and readiness endpoints](/operations/status-check) diff --git a/content/de/developer/integration/reverse-proxy/index.md b/content/de/developer/integration/reverse-proxy/index.md index eff7324d..397cdad6 100644 --- a/content/de/developer/integration/reverse-proxy/index.md +++ b/content/de/developer/integration/reverse-proxy/index.md @@ -13,6 +13,7 @@ We recommend using separate hostnames for the S3 API on port `9000` and the Cons - [Traefik](./traefik.md) - [Caddy](./caddy.md) - [HAProxy](./haproxy.md) +- [Envoy](./envoy.md) - [Apache HTTP Server](./httpd.md) ## Related configuration diff --git a/content/de/developer/integration/reverse-proxy/meta.json b/content/de/developer/integration/reverse-proxy/meta.json index adeb173e..47c8243f 100644 --- a/content/de/developer/integration/reverse-proxy/meta.json +++ b/content/de/developer/integration/reverse-proxy/meta.json @@ -4,7 +4,8 @@ "nginx", "traefik", "caddy", + "envoy", "haproxy", "httpd" ] -} \ No newline at end of file +} diff --git a/content/en/developer/integration/ai/images/rustfs-vllm-models.png b/content/en/developer/integration/ai/images/rustfs-vllm-models.png new file mode 100644 index 00000000..baa18dc1 Binary files /dev/null and b/content/en/developer/integration/ai/images/rustfs-vllm-models.png differ diff --git a/content/en/developer/integration/ai/index.md b/content/en/developer/integration/ai/index.md index 06cc8316..572f68bb 100644 --- a/content/en/developer/integration/ai/index.md +++ b/content/en/developer/integration/ai/index.md @@ -8,5 +8,6 @@ Use **RustFS** as the object storage layer for AI and machine learning platforms ## Platforms - [Ray](./ray.md) +- [vLLM](./vllm.md) Keep training datasets and checkpoints in dedicated buckets, and use credentials scoped to the required bucket operations. diff --git a/content/en/developer/integration/ai/meta.json b/content/en/developer/integration/ai/meta.json index 573fc520..b30f9e0f 100644 --- a/content/en/developer/integration/ai/meta.json +++ b/content/en/developer/integration/ai/meta.json @@ -1,6 +1,7 @@ { "title": "AI", "pages": [ - "ray" + "ray", + "vllm" ] } diff --git a/content/en/developer/integration/ai/vllm.md b/content/en/developer/integration/ai/vllm.md new file mode 100644 index 00000000..c471096c --- /dev/null +++ b/content/en/developer/integration/ai/vllm.md @@ -0,0 +1,154 @@ +--- +title: "vLLM" +description: "Serve LLM inference with vLLM loading model weights stored in RustFS." +--- + +This guide connects [vLLM](https://github.com/vllm-project/vllm) — the high-throughput LLM inference engine — to **RustFS** as its model-weight store. You will upload a model into a RustFS bucket, expose the bucket to the vLLM host through an rclone mount, and serve the model with the OpenAI-compatible API. The workflow was verified with `vllm/vllm-openai-cpu` (vLLM 0.30.0) serving `facebook/opt-125m` from a RustFS bucket backed by `rustfs/rustfs-x86-musl:v2.3.1`, on a CPU-only host. + +You need Docker and an rclone binary on the host that runs vLLM. This deployment is intended for local integration testing, not production. + +## Architecture + +```mermaid +flowchart LR + Upload["rclone copy"] -->|"weights"| RustFS["RustFS :9000"] + RustFS -->|"rclone mount"| Mount["/mnt/vllm-models"] + Mount -->|"weight load"| vLLM["vLLM :8000"] + Client["OpenAI SDK / curl"] -->|"completions"| vLLM +``` + +The bucket is the single copy of the model. Hosts that serve the model mount the bucket read-only, so every node pulls weights from RustFS and no local model store exists to drift. + +:::note[Why an rclone mount] + +vLLM 0.30 loads `s3://` model paths through the RunAI model streamer, whose ranged reads currently fail against custom S3 endpoints such as RustFS (the loader errors with `File access error` on any non-zero offset). Mounting the bucket as a filesystem is the verified way to keep the weights in RustFS while vLLM reads them as local files. + +::: + +## 1. Upload the model to RustFS + +Create the bucket and copy model weights into it, replacing all connection placeholders: + +```ini title="rclone.conf" +[rustfs] +type = s3 +provider = Other +access_key_id = +secret_access_key = +endpoint = http://:9000 +region = us-east-1 +``` + +```bash +rc mb rustfs/vllm-models +rclone copy ./opt-125m rustfs:vllm-models/opt-125m --transfers 4 +``` + +Any Hugging Face layout works — `config.json`, the tokenizer files, and the weight files (`model.safetensors` or `pytorch_model.bin`). Keep one model per prefix so several models can share the bucket. + +## 2. Mount the bucket on the vLLM host + +On the machine that runs vLLM, mount the bucket read-only for clients with `--allow-other`: + +```bash +mkdir -p /mnt/vllm-models +rclone mount rustfs:vllm-models /mnt/vllm-models \ + --allow-other --daemon +ls /mnt/vllm-models/opt-125m/ +``` + +```text +config.json merges.txt model.safetensors tokenizer.json vocab.json +``` + +## 3. Run vLLM + +Start the CPU image against the mounted weights: + +```bash +docker run -d --name vllm -p 8000:8000 --shm-size=2g \ + -v /mnt/vllm-models:/models:ro \ + vllm/vllm-openai-cpu:latest \ + --model /models/opt-125m --served-model-name opt-125m \ + --dtype float32 --max-model-len 256 --gpu-memory-utilization 0.15 +``` + +vLLM reads the weights through the mount — the container stays stateless and the model lives in RustFS. Wait for the server to come up: + +```bash +curl -s http://localhost:8000/v1/models | head -c 200 +``` + +```text +{"object":"list","data":[{"id":"opt-125m","object":"model","created":...,"root":"/models/opt-125m",...}]} +``` + +`--gpu-memory-utilization` controls the fraction of RAM reserved for the KV cache on the CPU backend; lower it on small hosts. `--dtype float32` matches what the CPU attention kernels support for this model. + +## 4. Run inference + +Send an OpenAI-compatible completion request: + +```bash +curl -s http://localhost:8000/v1/completions \ + -H "Content-Type: application/json" \ + -d '{"model": "opt-125m", "prompt": "RustFS is", "max_tokens": 12, "temperature": 0}' +``` + +```json +{"id":"cmpl-...","object":"text_completion","model":"opt-125m", + "choices":[{"index":0,"text":" a great tool for building your own server. It's a", + "finish_reason":"length",...}]} +``` + +The request is standard OpenAI schema, so the Python client works unchanged: + +```python +from openai import OpenAI + +client = OpenAI(base_url="http://localhost:8000/v1", api_key="EMPTY") +print(client.completions.create( + model="opt-125m", prompt="RustFS is", max_tokens=12, temperature=0, +).choices[0].text) +``` + +![vLLM model weights stored in the RustFS Console](./images/rustfs-vllm-models.png) + +## 5. Stop or reset + +To tear down the demo while keeping the bucket objects: + +```bash +docker rm -f vllm +fusermount -u /mnt/vllm-models +``` + +To delete the stored model: + +```bash +rclone purge rustfs:vllm-models +``` + +## Troubleshooting + +### `Cannot find any model weights with /models/...` + +The mount had a stale directory cache or the weight files never made it to the bucket. Re-run `rclone copy` and confirm the files through the mount with `ls` before starting vLLM. A short `--dir-cache-time` (for example `10s`) helps while you iterate. + +### `Unsupported CPU attention configuration: head_dim=...` + +vLLM's CPU kernels support a fixed set of head dimensions. Tiny test models such as `hf-internal-testing/tiny-random-*` use exotic shapes that fail at request time — use a real small model such as `facebook/opt-125m`. + +### `Insufficient space in /dev/shm` + +vLLM's CPU engine exchanges tensors through shared memory. Run the container with `--shm-size=2g` (or `--ipc=host`). + +### Server exits with `Available memory on node 0 ... is less than desired CPU memory utilization` + +The default KV-cache reservation is 90% of system RAM. Lower it with `--gpu-memory-utilization 0.15` (the flag applies to the CPU backend as a memory fraction despite its name). + +## Next steps + +- Review [S3 compatibility notes](/administration/protocols/s3) before adopting additional serving setups. +- Create dedicated production credentials with [Access Key Management](/security-compliance/iam/access-token). +- Follow the [vLLM documentation](https://docs.vllm.ai/en/latest/serving/openai_compatible_server.html) for chat templates, tensor parallelism, and quantized weights on top of the same bucket-backed model store. diff --git a/content/en/developer/integration/big-data/airflow.md b/content/en/developer/integration/big-data/airflow.md new file mode 100644 index 00000000..7f55e219 --- /dev/null +++ b/content/en/developer/integration/big-data/airflow.md @@ -0,0 +1,170 @@ +--- +title: "Airflow" +description: "Move data between Airflow DAGs and RustFS with the Amazon S3 provider." +--- + +This guide connects [Apache Airflow](https://github.com/apache/airflow) — the workflow orchestration platform — to **RustFS** through the Amazon S3 provider's hooks, operators, and sensors. You will register a custom-endpoint connection, run a DAG that writes an object to a RustFS bucket, waits for a key with `S3KeySensor`, and reads the object back with `S3Hook`. The workflow was verified with `apache/airflow:3.3.2` (standalone, SequentialExecutor) and the `apache-airflow-providers-amazon` provider against `rustfs/rustfs-x86-musl:v2.3.1`. + +You need Docker, or an existing Airflow installation. This deployment is intended for local integration testing, not production. + +## Architecture + +```mermaid +flowchart LR + Scheduler["Airflow scheduler"] -->|"tasks"| Hook["S3Hook / operators"] + Hook -->|"S3 API"| RustFS["RustFS :9000"] + Sensor["S3KeySensor"] -->|"poll key"| RustFS +``` + +Every S3 interaction inside a DAG goes through the provider's S3 client, pointed at RustFS by the connection's `endpoint_url`. Operators, sensors, and hooks share the same connection object. + +## 1. Run Airflow + +Start a standalone instance with examples disabled, and create the demo bucket: + +```bash +docker run -d --name airflow --network oo-rustfs_default -p 8080:8080 \ + -e AIRFLOW__CORE__LOAD_EXAMPLES=False \ + -v "$PWD/dags":/opt/airflow/dags \ + apache/airflow:3.3.2 standalone + +rc mb rustfs/airflow-demo +``` + +The image ships with all providers preinstalled, including `apache-airflow-providers-amazon`. + +## 2. Register the RustFS connection + +The S3 provider reads its endpoint from the connection's extra field. Replace all connection placeholders: + +```bash +docker exec airflow airflow connections add rustfs \ + --conn-type aws \ + --conn-extra '{"endpoint_url": "http://:9000", "region_name": "us-east-1", "aws_access_key_id": "", "aws_secret_access_key": ""}' +``` + +The keys `aws_access_key_id` and `aws_secret_access_key` inside `--conn-extra` supply credentials; `endpoint_url` redirects the boto3 client from AWS to RustFS. + +## 3. Write the DAG + +The DAG writes an object with an operator, waits for the key with a sensor, and reads it back with the hook: + +```python title="rustfs_demo.py" +import datetime + +from airflow.providers.amazon.aws.hooks.s3 import S3Hook +from airflow.providers.amazon.aws.operators.s3 import S3CreateObjectOperator +from airflow.providers.amazon.aws.sensors.s3 import S3KeySensor +from airflow.sdk import dag, task + +@dag( + schedule=None, + start_date=datetime.datetime(2026, 1, 1), + catchup=False, + tags=["rustfs"], +) +def rustfs_demo(): + create = S3CreateObjectOperator( + task_id="write_object", + s3_bucket="airflow-demo", + s3_key="dags/airflow-put.txt", + data="written by airflow to rustfs", + aws_conn_id="rustfs", + replace=True, + ) + + wait = S3KeySensor( + task_id="wait_for_object", + bucket_key="dags/airflow-put.txt", + bucket_name="airflow-demo", + aws_conn_id="rustfs", + timeout=120, + poke_interval=10, + mode="reschedule", + ) + + @task + def read_object(): + hook = S3Hook(aws_conn_id="rustfs") + body = hook.read_key(key="dags/airflow-put.txt", bucket_name="airflow-demo") + print("read back:", body) + assert body == "written by airflow to rustfs" + + create >> [wait, read_object()] + +rustfs_demo() +``` + +Note the import paths: `S3CreateObjectOperator` lives in the `operators` module while `S3KeySensor` lives in the `sensors` module — importing both from one place fails. + +## 4. Unpause and trigger + +New DAGs start paused, and a trigger fired while paused stays queued forever. Unpause first, then trigger: + +```bash +docker exec airflow airflow dags unpause rustfs_demo +docker exec airflow airflow dags trigger rustfs_demo +``` + +Watch the run finish: + +```bash +docker exec airflow airflow dags list-runs rustfs_demo | head -3 +``` + +```text +dag_id run_id state +rustfs_demo manual__2026-09-29T13:15:57.332688+00:00 success +``` + +All three tasks succeed: `write_object`, `wait_for_object`, and `read_object`. + +## 5. Verify objects in RustFS + +List the bucket prefix: + +```bash +rc ls rustfs/airflow-demo/ -r +rc cat rustfs/airflow-demo/dags/airflow-put.txt +``` + +```text +[2026-09-29 13:16:01] 28 B dags/airflow-put.txt +written by airflow to rustfs +``` + +![Airflow object stored in the RustFS Console](./images/rustfs-airflow-object.png) + +## 6. Stop or reset + +To tear down the demo while keeping the bucket objects: + +```bash +docker rm -f airflow +``` + +To delete the stored data: + +```bash +rc rm rustfs/airflow-demo/ --recursive --force +``` + +## Troubleshooting + +### Dag runs stay `queued` after triggering + +The DAG is paused. New DAGs are paused by default in Airflow 3, and runs triggered in that state never execute. Run `airflow dags unpause rustfs_demo`; queued runs then start on their own. + +### `cannot import name 'S3KeySensor' from 'airflow.providers.amazon.aws.operators.s3'` + +The sensor lives in a separate module: `from airflow.providers.amazon.aws.sensors.s3 import S3KeySensor`. + +### Tasks fail with connection errors + +The `endpoint_url` must be reachable from the Airflow container — use the Docker network hostname for RustFS, not `localhost`. Airflow 3 serves its health endpoint under `/api/v2/monitor/health` if you need to check component status. + +## Next steps + +- Review [S3 compatibility notes](/administration/protocols/s3) before adopting additional provider hooks. +- Create dedicated production credentials with [Access Key Management](/security-compliance/iam/access-token). +- Follow the [Amazon provider documentation](https://airflow.apache.org/docs/apache-airflow-providers-amazon/stable/index.html) for transfer operators such as `S3ToLocalFilesystemOperator` and `LocalFilesystemToS3Operator`. diff --git a/content/en/developer/integration/big-data/delta-lake.md b/content/en/developer/integration/big-data/delta-lake.md new file mode 100644 index 00000000..b1bfae56 --- /dev/null +++ b/content/en/developer/integration/big-data/delta-lake.md @@ -0,0 +1,125 @@ +--- +title: "Delta Lake" +description: "Write and read Delta tables on RustFS with delta-rs." +--- + +This guide connects [Delta Lake](https://github.com/delta-io/delta) — the open-source lakehouse table format — to **RustFS** through delta-rs, the Rust-native Delta implementation. You will write a Delta table to a RustFS bucket from Python, read it back with ACID transaction history, and confirm the `_delta_log` and Parquet files in the bucket. The workflow was verified with the `deltalake` Python package (delta-rs) and pandas against `rustfs/rustfs-x86-musl:v2.3.1`. + +You need Python 3.9 or newer. This deployment is intended for local integration testing, not production. + +## Architecture + +```mermaid +flowchart LR + DF["pandas DataFrame"] -->|"write_deltalake"| deltaRS["delta-rs"] + deltaRS -->|"Parquet + _delta_log"| RustFS["RustFS :9000"] + Query["DeltaTable"] -->|"read / time travel"| RustFS +``` + +delta-rs stores each table as Parquet files plus a transaction log (`_delta_log/`). All I/O goes through the `object_store` crate, configured with the same AWS environment variables as other S3 clients. + +## 1. Install the client + +```bash +pip install deltalake pandas pyarrow +``` + +`pyarrow` is required to convert pandas frames into Delta-compatible record batches. + +## 2. Write a Delta table + +Create the bucket and write a table, replacing all connection placeholders. `AWS_S3_ALLOW_UNSAFE_RENAME` is needed because RustFS does not provide copy-if-not-exists, which delta-rs otherwise uses for commit conflicts: + +```python title="delta_s3.py" +import pandas as pd +from deltalake import DeltaTable, write_deltalake + +storage_options = { + "AWS_ENDPOINT_URL": "http://:9000", + "AWS_ACCESS_KEY_ID": "", + "AWS_SECRET_ACCESS_KEY": "", + "AWS_REGION": "us-east-1", + "AWS_ALLOW_HTTP": "true", + "AWS_S3_ALLOW_UNSAFE_RENAME": "true", +} + +table = "s3:///events" +df = pd.DataFrame({"id": [1, 2, 3], "name": ["alpha", "beta", "gamma"]}) +write_deltalake(table, df, storage_options=storage_options) +print("written:", df.shape[0], "rows") +``` + +```text +written: 3 rows +``` + +The `table` URI uses the standard `s3://bucket/prefix` form; the endpoint and credentials come from `storage_options`. + +## 3. Read the table back + +```python title="delta_read.py" +from deltalake import DeltaTable + +back = DeltaTable("s3:///events", storage_options=storage_options).to_pandas() +print("read back:", back.shape[0], "rows") +print(back.sort_values("id").to_string(index=False)) +print("version:", DeltaTable("s3:///events", storage_options=storage_options).version()) +``` + +```text +read back: 3 rows + id name + 1 alpha + 2 beta + 3 gamma +version: 0 +``` + +Because the version is tracked in the transaction log, the same table supports time travel with `DeltaTable(..., version=N)` and appends that bump the version. + +## 4. Verify objects in RustFS + +List the table prefix: + +```bash +rc ls rustfs// -r +``` + +The first commit created the transaction log and one Parquet file: + +```text +events/_delta_log/00000000000000000000.json +events/part-00000-3859855e-45e4-4ae5-94ff-2d8eab5e7ebb-c000.snappy.parquet +``` + +Every new write adds a `NNNNNNNNNNNNNNNNNNNN.json` log entry and Parquet parts; readers replay the log to get a consistent snapshot. + +![Delta table files stored in the RustFS Console](./images/rustfs-delta-table.png) + +## 5. Stop or reset + +delta-rs holds no state of its own. To delete the table: + +```bash +rc rm rustfs//events/ --recursive --force +``` + +## Troubleshooting + +### `Import pyarrow failed` when writing a pandas DataFrame + +`write_deltalake` converts frames through Arrow. Install `pyarrow` alongside `deltalake` and `pandas`. + +### `Generic DeltaTable error: commit conflict` or rename errors on commit + +delta-rs commits by copying and renaming temporary objects, which requires atomic rename on the backend. For S3-compatible stores without copy-if-not-exists, set `AWS_S3_ALLOW_UNSAFE_RENAME: "true"` in `storage_options` — acceptable for a single writer, not for concurrent writers. + +### `Unknown lengthy error: AWS connectivity or endpoint errors` + +Confirm `AWS_ENDPOINT_URL` includes the scheme and that `AWS_ALLOW_HTTP` is `"true"` for plain-HTTP endpoints; without it the S3 client only speaks HTTPS. + +## Next steps + +- Review [S3 compatibility notes](/administration/protocols/s3) before adopting additional Delta clients. +- Create dedicated production credentials with [Access Key Management](/security-compliance/iam/access-token). +- Follow the [delta-rs usage documentation](https://delta-io.github.io/delta-rs/usage/writing/writing-to-s3/) for concurrent-writer setups with DynamoDB-backed commit coordination. diff --git a/content/en/developer/integration/big-data/images/rustfs-airflow-object.png b/content/en/developer/integration/big-data/images/rustfs-airflow-object.png new file mode 100644 index 00000000..9177d249 Binary files /dev/null and b/content/en/developer/integration/big-data/images/rustfs-airflow-object.png differ diff --git a/content/en/developer/integration/big-data/images/rustfs-delta-table.png b/content/en/developer/integration/big-data/images/rustfs-delta-table.png new file mode 100644 index 00000000..093a5cf1 Binary files /dev/null and b/content/en/developer/integration/big-data/images/rustfs-delta-table.png differ diff --git a/content/en/developer/integration/big-data/images/rustfs-kafka-sink.png b/content/en/developer/integration/big-data/images/rustfs-kafka-sink.png new file mode 100644 index 00000000..dd4b11b5 Binary files /dev/null and b/content/en/developer/integration/big-data/images/rustfs-kafka-sink.png differ diff --git a/content/en/developer/integration/big-data/index.md b/content/en/developer/integration/big-data/index.md index 78651eea..d26a8d96 100644 --- a/content/en/developer/integration/big-data/index.md +++ b/content/en/developer/integration/big-data/index.md @@ -8,6 +8,7 @@ Use **RustFS** as the object storage layer for data analytics systems that suppo ## Systems - [ClickHouse](./clickhouse.md) +- [Airflow](./airflow.md) - [Hudi](./hudi.md) - [Iceberg](./iceberg.md) - [PyIceberg](./pyiceberg.md) @@ -16,8 +17,10 @@ Use **RustFS** as the object storage layer for data analytics systems that suppo - [OpenDAL](./opendal.md) - [DuckDB](./duckdb.md) - [Doris](./doris.md) +- [Delta Lake](./delta-lake.md) - [lakeFS](./lakefs.md) - [InfluxDB](./influxdb.md) +- [Kafka](./kafka.md) - [Spark](./spark.md) - [Flink](./flink.md) - [Trino](./trino.md) diff --git a/content/en/developer/integration/big-data/kafka.md b/content/en/developer/integration/big-data/kafka.md new file mode 100644 index 00000000..ff1e92f9 --- /dev/null +++ b/content/en/developer/integration/big-data/kafka.md @@ -0,0 +1,178 @@ +--- +title: "Kafka" +description: "Offload Kafka topic data to RustFS with the Kafka Connect S3 sink connector." +--- + +This guide connects [Apache Kafka](https://github.com/apache/kafka) — the distributed event streaming platform — to **RustFS** through the Kafka Connect S3 sink connector. You will run a KRaft broker and a Connect worker, deploy the S3 sink for a topic, and produce records that land as objects in a RustFS bucket. The workflow was verified with `apache/kafka:4.0.0` and `confluentinc/kafka-connect-s3` v10.5.25 against `rustfs/rustfs-x86-musl:v2.3.1`. + +Kafka's KIP-405 tiered storage needs a `RemoteLogStorageManager` plugin, and the S3 implementations in the ecosystem are vendor-proprietary. The Connect S3 sink is the open, self-hosted way to move topic data to S3-compatible storage and is the approach documented here. + +You need Docker. This deployment is intended for local integration testing, not production. + +## Architecture + +```mermaid +flowchart LR + Producer["Console producer"] -->|"records"| Broker["Kafka broker :9092"] + Broker -->|"consumer group"| Connect["Connect S3 sink"] + Connect -->|"batched objects"| RustFS["RustFS :9000"] +``` + +The sink task consumes a topic in a dedicated consumer group and writes record batches to the bucket, one object per `flush.size` records per partition. + +## 1. Run the broker + +Start a KRaft broker whose advertised listener is reachable from other containers: + +```bash +docker run -d --name kafka --hostname kafka --network oo-rustfs_default \ + -e CLUSTER_ID=5L6g3nShT-eMCtK--X86sw \ + -e KAFKA_NODE_ID=1 \ + -e KAFKA_PROCESS_ROLES=broker,controller \ + -e KAFKA_LISTENERS=PLAINTEXT://:9092,CONTROLLER://:9093 \ + -e KAFKA_ADVERTISED_LISTENERS=PLAINTEXT://kafka:9092 \ + -e KAFKA_CONTROLLER_LISTENER_NAMES=CONTROLLER \ + -e KAFKA_LISTENER_SECURITY_PROTOCOL_MAP=CONTROLLER:PLAINTEXT,PLAINTEXT:PLAINTEXT \ + -e KAFKA_CONTROLLER_QUORUM_VOTERS=1@kafka:9093 \ + -e KAFKA_OFFSETS_TOPIC_REPLICATION_FACTOR=1 \ + -e KAFKA_TRANSACTION_STATE_LOG_REPLICATION_FACTOR=1 \ + -e KAFKA_TRANSACTION_STATE_LOG_MIN_ISR=1 \ + apache/kafka:4.0.0 + +docker exec kafka /opt/kafka/bin/kafka-topics.sh \ + --bootstrap-server localhost:9092 \ + --create --topic rustfs-topic --partitions 1 --replication-factor 1 + +rc mb rustfs/kafka-demo +``` + +The image defaults to advertising `localhost:9092`, which only works from inside the broker container. The `KAFKA_ADVERTISED_LISTENERS` override is what makes Connect (and any remote client) able to reach the broker. + +## 2. Install the connector + +Download the Confluent Hub archive, which bundles the connector and its dependencies, and unpack it where the worker can see it: + +```bash +curl -Lo kafka-connect-s3.zip "https://hub-downloads.confluent.io/api/plugins/confluentinc/kafka-connect-s3/versions/10.5.25/confluentinc-kafka-connect-s3-10.5.25.zip" +unzip kafka-connect-s3.zip -d /opt/kafka-conn/plugins +``` + +## 3. Configure the worker and the sink + +Create the worker properties. The value converter must be `ByteArrayConverter` so records are written verbatim: + +```ini title="worker.properties" +bootstrap.servers=kafka:9092 +key.converter=org.apache.kafka.connect.storage.StringConverter +value.converter=org.apache.kafka.connect.converters.ByteArrayConverter +offset.storage.file.filename=/tmp/connect.offsets +offset.flush.interval.ms=5000 +plugin.path=/opt/kafka-conn-plugins +``` + +Create the sink connector configuration, replacing all connection placeholders: + +```ini title="rustfs-sink.properties" +name=rustfs-sink +connector.class=io.confluent.connect.s3.S3SinkConnector +tasks.max=1 +topics=rustfs-topic +s3.bucket.name=kafka-demo +s3.region=us-east-1 +store.url=http://:9000 +s3.path.style.access.enabled=true +flush.size=3 +storage.class=io.confluent.connect.s3.storage.S3Storage +format.class=io.confluent.connect.s3.format.bytearray.ByteArrayFormat +consumer.override.auto.offset.reset=earliest +``` + +Kafka 4.0 moved the class to `org.apache.kafka.connect.converters.ByteArrayConverter` — the old `storage` package path no longer resolves. + +## 4. Run the worker + +Run `connect-standalone` in the foreground so its logs go to `docker logs`, with the bucket credentials in the environment: + +```bash +docker run -d --name kafka-connect --hostname kafka-connect \ + --network oo-rustfs_default \ + -v /opt/kafka-conn/plugins:/opt/kafka-conn-plugins:ro \ + -v "$PWD/worker.properties":/etc/kafka/worker.properties:ro \ + -v "$PWD/rustfs-sink.properties":/etc/kafka/sink.properties:ro \ + -e AWS_ACCESS_KEY_ID= \ + -e AWS_SECRET_ACCESS_KEY= \ + apache/kafka:4.0.0 \ + /opt/kafka/bin/connect-standalone.sh /etc/kafka/worker.properties /etc/kafka/sink.properties +``` + +The worker is ready when the sink task claims the partition: + +```text +INFO [rustfs-sink|task-0] Assigned topic partitions: [rustfs-topic-0] +``` + +## 5. Produce records + +Send at least `flush.size` records so the connector completes a batch: + +```bash +docker exec kafka sh -c "printf 'msg-one\nmsg-two\nmsg-three\n' | \ + /opt/kafka/bin/kafka-console-producer.sh --bootstrap-server localhost:9092 --topic rustfs-topic" +``` + +After a few seconds the batch becomes an object in the bucket: + +```bash +rc ls rustfs/kafka-demo/ -r +rc cat rustfs/kafka-demo/topics/rustfs-topic/partition=0/rustfs-topic+0+0000000000.bin +``` + +```text +topics/rustfs-topic/partition=0/rustfs-topic+0+0000000000.bin +topics/rustfs-topic/partition=0/rustfs-topic+0+0000000003.bin +msg-one +msg-two +msg-three +``` + +The object name encodes topic, partition, and starting offset. Each subsequent batch of three records lands in the next object (`+0000000003.bin` and so on). + +![Kafka sink objects stored in the RustFS Console](./images/rustfs-kafka-sink.png) + +## 6. Stop or reset + +To tear down the demo while keeping the bucket objects: + +```bash +docker rm -f kafka-connect kafka +``` + +To delete the stored data: + +```bash +rc rm rustfs/kafka-demo/ --recursive --force +``` + +## Troubleshooting + +### `AdminClient ... Rebootstrapping with Cluster (id: null)` loops forever + +The broker advertises `localhost:9092`, so a remote client receives metadata pointing at itself. Set `KAFKA_ADVERTISED_LISTENERS=PLAINTEXT://kafka:9092` (and matching listener variables) as in step 1. + +### `Invalid schema type for ByteArrayConverter: STRING` + +The `ByteArrayFormat` writer only accepts raw bytes. Either switch the worker's `value.converter` to the ByteArray converter or choose a format class that matches the converter you use. + +### `Class org.apache.kafka.connect.storage.ByteArrayConverter could not be found` + +Kafka 4.0 moved the class to `org.apache.kafka.connect.converters.ByteArrayConverter`. Use the new package path in `worker.properties`. + +### The connector downloads but the plugin is not found + +The plain connector JAR from Maven lacks its dependencies. Use the Confluent Hub archive from step 2, which bundles the complete `lib/` directory, and make sure `plugin.path` points at the directory that contains the connector folder. + +## Next steps + +- Review [S3 compatibility notes](/administration/protocols/s3) before adopting additional connectors. +- Create dedicated production credentials with [Access Key Management](/security-compliance/iam/access-token). +- Follow the [Kafka Connect S3 sink documentation](https://docs.confluent.io/kafka-connect-s3/current/index.html) for Parquet and Avro formats, partitioning by time, and IAM-based credential chains. diff --git a/content/en/developer/integration/big-data/meta.json b/content/en/developer/integration/big-data/meta.json index 29684232..2374fa35 100644 --- a/content/en/developer/integration/big-data/meta.json +++ b/content/en/developer/integration/big-data/meta.json @@ -2,12 +2,15 @@ "title": "Data Analytics", "pages": [ "clickhouse", + "airflow", "duckdb", "doris", + "delta-lake", "flink", "hudi", "iceberg", "influxdb", + "kafka", "lakefs", "milvus", "mlflow", diff --git a/content/en/developer/integration/devops/images/rustfs-opensearch-snapshot.png b/content/en/developer/integration/devops/images/rustfs-opensearch-snapshot.png new file mode 100644 index 00000000..60b3bd3a Binary files /dev/null and b/content/en/developer/integration/devops/images/rustfs-opensearch-snapshot.png differ diff --git a/content/en/developer/integration/devops/index.md b/content/en/developer/integration/devops/index.md index ad31918d..a3c404a2 100644 --- a/content/en/developer/integration/devops/index.md +++ b/content/en/developer/integration/devops/index.md @@ -8,6 +8,7 @@ Use **RustFS** as the object storage layer for DevOps platforms and infrastructu ## Platforms - [Elasticsearch](./elasticsearch.md) +- [OpenSearch](./opensearch.md) - [Gitea](./gitea.md) - [Jenkins](./jenkins.md) - [Terraform](./terraform.md) diff --git a/content/en/developer/integration/devops/meta.json b/content/en/developer/integration/devops/meta.json index 67c84c6e..55eada6a 100644 --- a/content/en/developer/integration/devops/meta.json +++ b/content/en/developer/integration/devops/meta.json @@ -2,6 +2,7 @@ "title": "DevOps", "pages": [ "elasticsearch", + "opensearch", "gitea", "jenkins", "terraform" diff --git a/content/en/developer/integration/devops/opensearch.md b/content/en/developer/integration/devops/opensearch.md new file mode 100644 index 00000000..9d44e530 --- /dev/null +++ b/content/en/developer/integration/devops/opensearch.md @@ -0,0 +1,172 @@ +--- +title: "OpenSearch" +description: "Snapshot OpenSearch indices to RustFS with the repository-s3 plugin." +--- + +This guide connects [OpenSearch](https://github.com/opensearch-project/OpenSearch) — the open-source search and analytics suite derived from Elasticsearch — to **RustFS** through the `repository-s3` plugin. You will register an S3 snapshot repository backed by a RustFS bucket, take a snapshot of an index, and restore it. The workflow was verified with `opensearchproject/opensearch:3.8.0` and the bundled `repository-s3` plugin against `rustfs/rustfs-x86-musl:v2.3.1`. + +You need Docker, or an OpenSearch node where you can install plugins and edit configuration. This deployment is intended for local integration testing, not production. + +## Architecture + +```mermaid +flowchart LR + Client["REST client"] --> OS["OpenSearch :9200"] + OS -->|"snapshot files"| RustFS["RustFS :9000"] + RustFS -->|"restore"| OS +``` + +The `repository-s3` plugin writes snapshots as shard archives plus metadata blobs in the bucket. Registration is cluster-wide, so every node needs the plugin and the same client configuration. + +## 1. Run OpenSearch + +Start a single node with security disabled and a small heap: + +```bash +docker run -d --name opensearch --network oo-rustfs_default -p 9200:9200 \ + -e discovery.type=single-node \ + -e OPENSEARCH_JAVA_OPTS="-Xms512m -Xmx512m" \ + -e DISABLE_SECURITY_PLUGIN=true \ + opensearchproject/opensearch:3.8.0 +``` + +The node is ready when `curl http://localhost:9200` returns the cluster header (allow one to two minutes). + +## 2. Install the repository-s3 plugin + +The S3 repository plugin is not preloaded. Install it and restart the node: + +```bash +docker exec opensearch bin/opensearch-plugin install --batch repository-s3 +docker restart opensearch +``` + +## 3. Configure the S3 client + +Credentials are secure settings: they belong in the OpenSearch keystore, not in the repository request or `opensearch.yml`. Create the keystore entries, replacing all connection placeholders: + +```bash +docker exec opensearch sh -c \ + "printf '' | bin/opensearch-keystore create 2>/dev/null; \ + printf '' | bin/opensearch-keystore add -f -x s3.client.default.access_key; \ + printf '' | bin/opensearch-keystore add -f -x s3.client.default.secret_key" +``` + +Add the non-secure client settings to `config/opensearch.yml`: + +```yaml title="opensearch.yml" +network.host: 0.0.0.0 +plugins.security.disabled: true +s3.client.default.endpoint: http://:9000 +s3.client.default.protocol: http +s3.client.default.path_style_access: "true" +``` + +Restart the node once more so it reads both the keystore and the new settings: + +```bash +docker restart opensearch +``` + +Create the bucket while the node boots: + +```bash +rc mb rustfs/opensearch-snapshots +``` + +## 4. Register the repository and snapshot + +Create a test index with a document, then register the repository: + +```bash +curl -sX PUT http://localhost:9200/rustfs-demo -H "Content-Type: application/json" \ + -d '{"settings":{"number_of_shards":1}}' + +curl -sX PUT http://localhost:9200/rustfs-demo/_doc/1 -H "Content-Type: application/json" \ + -d '{"product":"rustfs","via":"opensearch-snapshot"}' + +curl -sX PUT "http://localhost:9200/_snapshot/rustfs-repo" -H "Content-Type: application/json" \ + -d '{"type":"s3","settings":{"bucket":"opensearch-snapshots","region":"us-east-1","server_side_encryption_type":"bucket_default"}}' +``` + +The `server_side_encryption_type: bucket_default` setting matters: without it the plugin requests SSE-S3, which a self-hosted RustFS without a server-side encryption master key rejects. + +Take a snapshot and wait for completion: + +```bash +curl -sX PUT "http://localhost:9200/_snapshot/rustfs-repo/snapshot-1?wait_for_completion=true" \ + -H "Content-Type: application/json" -d '{"indices":"rustfs-demo"}' +``` + +```text +{"snapshot":{"snapshot":"snapshot-1","state":"SUCCESS","indices":["rustfs-demo"],...}} +``` + +## 5. Verify objects and restore + +List the bucket: + +```bash +rc ls rustfs/opensearch-snapshots/ -r +``` + +```text +index-0 +index.latest +indices/5x1bwsWaSv2XINIbeoe-RQ/0/__GgxvoCw-TKuMMAYBq5Khag +indices/5x1bwsWaSv2XINIbeoe-RQ/0/snap-kRgFBiuyQIyPo3_-CMp7Hw.dat +meta-kRgFBiuyQIyPo3_-CMp7Hw.dat +snap-kRgFBiuyQIyPo3_-CMp7Hw.dat +``` + +Delete the index and restore it from the snapshot: + +```bash +curl -sX DELETE http://localhost:9200/rustfs-demo +curl -sX POST "http://localhost:9200/_snapshot/rustfs-repo/snapshot-1/_restore?wait_for_completion=true" +curl -s http://localhost:9200/rustfs-demo/_doc/1 +``` + +```text +{"_index":"rustfs-demo","_id":"1","found":true,"_source":{"product":"rustfs","via":"opensearch-snapshot"}} +``` + +![OpenSearch snapshot stored in the RustFS Console](./images/rustfs-opensearch-snapshot.png) + +## 6. Stop or reset + +To tear down the demo while keeping the bucket objects: + +```bash +docker rm -f opensearch +``` + +To delete the stored snapshots: + +```bash +rc rm rustfs/opensearch-snapshots/ --recursive --force +``` + +## Troubleshooting + +### `Setting [access_key] is insecure, but property [allow_insecure_settings] is not set` + +Inline credentials in the repository request are rejected. Store them in the keystore as shown in step 3 — `access_key` and `secret_key` are secure settings in OpenSearch. + +### `SSE-S3 requires RUSTFS_SSE_S3_MASTER_KEY ... (Status Code: 400)` + +The plugin encrypts uploads with SSE-S3 by default. Register the repository with `"server_side_encryption_type": "bucket_default"` so no encryption header is sent, as in step 4. + +### `unknown setting [s3.client.default.access_key]` at startup + +The settings reference the repository-s3 plugin. If the node fails to start with them present, the plugin is not installed in that container — repeat step 2 (a fresh container loses plugins installed with `docker exec`). + +### Repository verification fails with `path is not accessible` + +The node cannot reach the bucket: check that `s3.client.default.endpoint` is reachable from the container, `path_style_access` is `"true"`, and the keystore credentials were loaded (they are read at startup — restart after adding them). + +## Next steps + +- Review [S3 compatibility notes](/administration/protocols/s3) before adopting additional OpenSearch repositories. +- Create dedicated production credentials with [Access Key Management](/security-compliance/iam/access-token). +- Follow the [OpenSearch snapshots documentation](https://docs.opensearch.org/docs/latest/tuning-your-cluster/availability-and-recovery/snapshots/index/) to automate snapshots with Snapshot Management (SM) policies. diff --git a/content/en/developer/integration/index.md b/content/en/developer/integration/index.md index 12798bd1..6957a392 100644 --- a/content/en/developer/integration/index.md +++ b/content/en/developer/integration/index.md @@ -7,14 +7,14 @@ Use this section to connect **RustFS** to infrastructure and application platfor ## Integration categories -- [Reverse Proxy](./reverse-proxy/index.md) covers Nginx, Traefik, Caddy, and HAProxy. +- [Reverse Proxy](./reverse-proxy/index.md) covers Nginx, Traefik, Caddy, HAProxy, and Envoy. - [Backup](./backup/index.md) covers Kopia, Longhorn, Restic, and Velero. -- [AI](./ai/index.md) covers AI platforms including Ray. -- [Data Analytics](./big-data/index.md) covers analytics systems including ClickHouse, Doris, Hudi, Iceberg, lakeFS, Milvus, OpenDAL, Vitess, and Zeppelin. +- [AI](./ai/index.md) covers AI platforms including Ray and vLLM. +- [Data Analytics](./big-data/index.md) covers analytics systems including Airflow, ClickHouse, Delta Lake, Doris, Hudi, Iceberg, Kafka, lakeFS, Milvus, OpenDAL, Vitess, and Zeppelin. - [Cloud Native](./cloud-native/index.md) covers Cortex and Flux. -- [Observability](./observability/index.md) covers telemetry systems including Fluentd, OpenObserve, OpenTelemetry, Thanos, and Tempo. -- [Others](./others/index.md) covers the community-driven capo SDK for Python. +- [Observability](./observability/index.md) covers telemetry systems including Fluentd, GreptimeDB, Loki, OpenObserve, OpenTelemetry, Tempo, Thanos, and VictoriaMetrics. +- [Others](./others/index.md) covers the capo SDK, rclone, JuiceFS, Nextcloud, and tusd. - [Registry](./registry/index.md) covers Harbor. -- [DevOps](./devops/index.md) covers Elasticsearch, Gitea, Jenkins, and Terraform. +- [DevOps](./devops/index.md) covers Elasticsearch, Gitea, Jenkins, OpenSearch, and Terraform. Each guide identifies the RustFS endpoint and addressing requirements to use when configuring the integrating system. \ No newline at end of file diff --git a/content/en/developer/integration/observability/greptimedb.md b/content/en/developer/integration/observability/greptimedb.md new file mode 100644 index 00000000..5b7ee690 --- /dev/null +++ b/content/en/developer/integration/observability/greptimedb.md @@ -0,0 +1,129 @@ +--- +title: "GreptimeDB" +description: "Run GreptimeDB with RustFS as the S3-compatible object storage backend." +--- + +This guide connects [GreptimeDB](https://github.com/GreptimeTeam/greptimedb) — the open-source, cloud-native time-series database — to **RustFS** as its object storage backend. You will start a standalone instance with its `[storage]` section pointed at a RustFS bucket, write time-series rows through the SQL API, and confirm the Parquet files and manifests in the bucket. The workflow was verified with `greptime/greptimedb` (main, commit `179ff8e5`) against `rustfs/rustfs-x86-musl:v2.3.1`. + +You need Docker, or a local GreptimeDB binary. This deployment is intended for local integration testing, not production. + +## Architecture + +```mermaid +flowchart LR + SQL["SQL / Prometheus API"] --> DB["GreptimeDB"] + DB -->|"SST + manifests"| RustFS["RustFS :9000"] +``` + +GreptimeDB keeps its write-ahead log and recent data locally, then persists SSTables (Parquet) and table manifests to object storage. Pointing the storage backend at RustFS makes the bucket the durable home of all table data. + +## 1. Configure the storage backend + +Create the bucket and a config file with an S3 storage section, replacing all connection placeholders: + +```toml title="greptimedb.toml" +[storage] +type = "S3" +bucket = "" +root = "greptimedb" +access_key_id = "" +secret_access_key = "" +endpoint = "http://:9000" +region = "us-east-1" +``` + +GreptimeDB uses path-style requests for custom endpoints by default; virtual-hosted style must be opted into explicitly with `enable_virtual_host_style`, so no extra flag is needed for RustFS. + +## 2. Run GreptimeDB + +Start a standalone instance with the config file: + +```bash +docker run -d --name greptimedb --network oo-rustfs_default -p 4000:4000 -p 4002:4002 \ + -v "$PWD/greptimedb.toml":/etc/greptimedb/greptimedb.toml:ro \ + greptime/greptimedb:latest standalone start \ + --http-addr 0.0.0.0:4000 \ + --mysql-addr 0.0.0.0:4002 \ + --config-file /etc/greptimedb/greptimedb.toml +``` + +Port `4000` serves the HTTP SQL endpoint and `4002` the MySQL protocol. + +## 3. Write and query time series + +Create a table, insert rows, and read them back. The HTTP SQL endpoint takes form-encoded requests: + +```bash +curl -s -X POST "http://localhost:4000/v1/sql" \ + --data-urlencode "sql=CREATE TABLE rustfs_demo (host STRING, cpu DOUBLE, mem DOUBLE, ts TIMESTAMP TIME INDEX)" + +curl -s -X POST "http://localhost:4000/v1/sql" \ + --data-urlencode "sql=INSERT INTO rustfs_demo VALUES (\"node-1\", 0.31, 0.62, 1790681000000), (\"node-1\", 0.35, 0.63, 1790681060000), (\"node-2\", 0.51, 0.71, 1790681000000)" +``` + +```text +{"output":[{"affectedrows":3}],"execution_time_ms":2} +``` + +Query the rows back: + +```bash +curl -s -X POST "http://localhost:4000/v1/sql" \ + --data-urlencode "sql=SELECT * FROM rustfs_demo ORDER BY ts" +``` + +```text +{"output":[{"records":{"rows":[["node-2",0.51,0.71,1790681000000],["node-1",0.35,0.63,1790681060000]],"total_rows":2}}]} +``` + +## 4. Verify objects in RustFS + +List the bucket — after the memtable flushes, the bucket holds Parquet SSTables and JSON manifests: + +```bash +rc ls rustfs// -r +``` + +```text +greptimedb/data/greptime/public/1024/1024_0000000000/manifest/00000000000000000000.json +greptimedb/data/greptime/greptime_private/1025/1025_0000000000/b11e8b25-5763-4f05-bcab-b6ee0a756a69.parquet +greptimedb/data/greptime/greptime_private/1025/1025_0000000000/manifest/00000000000000000001.json +``` + +Each database gets a directory under `data/`, and per-region `manifest/*.json` files describe the SSTables GreptimeDB reads back during queries. + +![GreptimeDB data stored in the RustFS Console](./images/rustfs-greptimedb-data.png) + +## 5. Stop or reset + +To tear down the demo while keeping the bucket objects: + +```bash +docker rm -f greptimedb +``` + +To delete the stored data: + +```bash +rc rm rustfs// --recursive --force +``` + +## Troubleshooting + +### `Form requests must have Content-Type: application/x-www-form-urlencoded` + +The `/v1/sql` HTTP endpoint only accepts form-encoded bodies. Pass SQL with `curl --data-urlencode "sql=..."` (or `application/x-www-form-urlencoded`), not as a JSON body. + +### Bucket stays empty + +GreptimeDB flushes memtables to object storage asynchronously. Run a few more inserts and wait a few seconds, or trigger a manual flush, then list the bucket again. + +### Startup fails with an S3 error + +Confirm `endpoint` includes the scheme, the bucket exists, and `access_key_id`/`secret_access_key` match a RustFS access key. The `root` value is optional but keeps the table tree under a known prefix. + +## Next steps + +- Review [S3 compatibility notes](/administration/protocols/s3) before adopting additional GreptimeDB storage options. +- Create dedicated production credentials with [Access Key Management](/security-compliance/iam/access-token). +- Follow the [GreptimeDB configuration reference](https://docs.greptime.com/operational-guide/configure/configure-datanode/) to tune flush intervals and cache layers for production workloads. diff --git a/content/en/developer/integration/observability/images/rustfs-greptimedb-data.png b/content/en/developer/integration/observability/images/rustfs-greptimedb-data.png new file mode 100644 index 00000000..1874f210 Binary files /dev/null and b/content/en/developer/integration/observability/images/rustfs-greptimedb-data.png differ diff --git a/content/en/developer/integration/observability/images/rustfs-vm-backups.png b/content/en/developer/integration/observability/images/rustfs-vm-backups.png new file mode 100644 index 00000000..a0bd0cee Binary files /dev/null and b/content/en/developer/integration/observability/images/rustfs-vm-backups.png differ diff --git a/content/en/developer/integration/observability/index.md b/content/en/developer/integration/observability/index.md index dbce4d0c..86e25a28 100644 --- a/content/en/developer/integration/observability/index.md +++ b/content/en/developer/integration/observability/index.md @@ -8,10 +8,12 @@ Use **RustFS** as the object storage layer for observability platforms that supp ## Platforms - [Fluentd](./fluentd.md) +- [GreptimeDB](./greptimedb.md) - [OpenObserve](./openobserve.md) - [OpenTelemetry](./opentelemetry.md) - [Loki](./loki.md) - [Tempo](./tempo.md) - [Thanos](./thanos.md) +- [VictoriaMetrics](./victoriametrics.md) Keep telemetry data in a dedicated bucket, and use credentials scoped to the required bucket operations. diff --git a/content/en/developer/integration/observability/meta.json b/content/en/developer/integration/observability/meta.json index d6ddebd7..97ef4d6e 100644 --- a/content/en/developer/integration/observability/meta.json +++ b/content/en/developer/integration/observability/meta.json @@ -2,10 +2,12 @@ "title": "Observability", "pages": [ "fluentd", + "greptimedb", "loki", "openobserve", "opentelemetry", "tempo", - "thanos" + "thanos", + "victoriametrics" ] } diff --git a/content/en/developer/integration/observability/victoriametrics.md b/content/en/developer/integration/observability/victoriametrics.md new file mode 100644 index 00000000..efc799d9 --- /dev/null +++ b/content/en/developer/integration/observability/victoriametrics.md @@ -0,0 +1,157 @@ +--- +title: "VictoriaMetrics" +description: "Back up VictoriaMetrics snapshots to RustFS with vmbackup." +--- + +This guide connects [VictoriaMetrics](https://github.com/VictoriaMetrics/VictoriaMetrics) — the Prometheus-compatible time-series database — to **RustFS** through `vmbackup` and `vmrestore`. You will run a single-node instance, import metrics, create an instant snapshot, back it up to a RustFS bucket, and restore the data into a fresh directory. The workflow was verified with `victoria-metrics`, `vmbackup`, and `vmrestore` v1.x images against `rustfs/rustfs-x86-musl:v2.3.1`. + +You need Docker. This deployment is intended for local integration testing, not production. + +## Architecture + +```mermaid +flowchart LR + Import["Prometheus import API"] --> VM["VictoriaMetrics :8428"] + VM -->|"instant snapshot"| Backup["vmbackup"] + Backup -->|"S3 API"| RustFS["RustFS :9000"] + RustFS -->|"restore"| Restore["vmrestore"] +``` + +`vmbackup` uploads a consistent point-in-time snapshot of the storage directory to any S3-compatible endpoint. `vmrestore` reverses the process, producing a data directory a VictoriaMetrics instance can open directly. + +## 1. Run VictoriaMetrics + +Create the bucket and start a single-node instance: + +```bash +rc mb rustfs/vm-backups + +docker run -d --name vm --network oo-rustfs_default -p 8428:8428 \ + -v vm-data:/storage \ + victoriametrics/victoria-metrics:latest \ + -storageDataPath=/storage -retentionPeriod=100y +``` + +## 2. Import metrics + +Write a couple of samples through the Prometheus import API: + +```bash +echo "vm_demo_metric 123" | curl -s --data-binary @- http://localhost:8428/api/v1/import/prometheus +echo "vm_demo_metric 456" | curl -s --data-binary @- http://localhost:8428/api/v1/import/prometheus +``` + +The endpoint answers `204 No Content`. Confirm the data is queryable: + +```bash +curl -s "http://localhost:8428/api/v1/export?match[]=vm_demo_metric" +``` + +```text +{"metric":{"__name__":"vm_demo_metric"},"values":[123,456],"timestamps":[1790680894604,1790680894619]} +``` + +## 3. Create a snapshot + +Ask VictoriaMetrics for a consistent snapshot: + +```bash +curl -s http://localhost:8428/snapshot/create +``` + +```text +{"status":"ok","snapshot":"20260929112134-18D9C6C76704F913"} +``` + +## 4. Back the snapshot up to RustFS + +Run `vmbackup` against the same storage volume, replacing the credential placeholders. The snapshot name comes from step 3: + +```bash +docker run --rm --network oo-rustfs_default \ + -e AWS_ACCESS_KEY_ID= \ + -e AWS_SECRET_ACCESS_KEY= \ + --volumes-from vm \ + victoriametrics/vmbackup:latest \ + -storageDataPath=/storage \ + -snapshotName=20260929112134-18D9C6C76704F913 \ + -dst=s3://vm-backups/demo \ + -customS3Endpoint=http://:9000 +``` + +```text +backup ... to S3{bucket: "vm-backups", dir: "demo/"} is complete; uploaded 760 bytes +``` + +`-customS3Endpoint` redirects the AWS SDK to RustFS; custom endpoints are addressed with path-style requests automatically. Set `AWS_EC2_METADATA_DISABLED=true` on hosts without an EC2 metadata service to skip credential lookup delays. + +## 5. Verify and restore + +List the bucket prefix: + +```bash +rc ls rustfs/vm-backups/demo/ +``` + +```text +backup_complete.ignore +backup_metadata.ignore +data/ +metadata/ +``` + +`backup_complete.ignore` marks a complete backup. Restore it into a fresh directory: + +```bash +docker run --rm --network oo-rustfs_default \ + -e AWS_ACCESS_KEY_ID= \ + -e AWS_SECRET_ACCESS_KEY= \ + -v /opt/vm-restore:/restore \ + victoriametrics/vmrestore:latest \ + -src=s3://vm-backups/demo \ + -storageDataPath=/restore \ + -customS3Endpoint=http://:9000 +``` + +```text +restored 760 bytes from backup in 0.055 seconds +``` + +The restored directory contains `data/`, `metadata/`, and a lock file — exactly what a VictoriaMetrics instance expects at `-storageDataPath`. + +![VictoriaMetrics backup stored in the RustFS Console](./images/rustfs-vm-backups.png) + +## 6. Stop or reset + +To tear down the demo while keeping the bucket objects: + +```bash +docker rm -f vm +docker volume rm vm-data +``` + +To delete the stored backups: + +```bash +rc rm rustfs/vm-backups/ --recursive --force +``` + +## Troubleshooting + +### `vmbackup` hangs at startup or fails to find credentials + +The AWS SDK probes the EC2 metadata service when environment credentials are absent. On machines without IMDS, export `AWS_EC2_METADATA_DISABLED=true` next to the key variables. + +### Backup parts re-upload on every run + +`vmbackup` performs incremental backups by comparing local and remote file hashes. Restoring to a fresh directory and running `vmbackup` from there re-uploads everything; keep the original data directory for incremental runs. + +### Query returns nothing right after import + +Imports are accepted asynchronously and the instant query endpoint can lag on a busy single node. Verify with `/api/v1/export` (or wait a few seconds) before creating the snapshot. + +## Next steps + +- Review [S3 compatibility notes](/administration/protocols/s3) before adopting additional VictoriaMetrics components. +- Create dedicated production credentials with [Access Key Management](/security-compliance/iam/access-token). +- Follow the [vmbackup documentation](https://docs.victoriametrics.com/vmbackup/) to schedule backups and prune old snapshots. diff --git a/content/en/developer/integration/others/images/rustfs-juicefs-chunks.png b/content/en/developer/integration/others/images/rustfs-juicefs-chunks.png new file mode 100644 index 00000000..737490b6 Binary files /dev/null and b/content/en/developer/integration/others/images/rustfs-juicefs-chunks.png differ diff --git a/content/en/developer/integration/others/images/rustfs-nextcloud-file.png b/content/en/developer/integration/others/images/rustfs-nextcloud-file.png new file mode 100644 index 00000000..506e0d79 Binary files /dev/null and b/content/en/developer/integration/others/images/rustfs-nextcloud-file.png differ diff --git a/content/en/developer/integration/others/images/rustfs-rclone-sync.png b/content/en/developer/integration/others/images/rustfs-rclone-sync.png new file mode 100644 index 00000000..dcf13712 Binary files /dev/null and b/content/en/developer/integration/others/images/rustfs-rclone-sync.png differ diff --git a/content/en/developer/integration/others/images/rustfs-tus-uploads.png b/content/en/developer/integration/others/images/rustfs-tus-uploads.png new file mode 100644 index 00000000..e82a55e1 Binary files /dev/null and b/content/en/developer/integration/others/images/rustfs-tus-uploads.png differ diff --git a/content/en/developer/integration/others/index.md b/content/en/developer/integration/others/index.md index 2a042b3e..65a45a2d 100644 --- a/content/en/developer/integration/others/index.md +++ b/content/en/developer/integration/others/index.md @@ -8,3 +8,7 @@ Integration guides that do not fit the other categories. ## Guides - [capo (Python)](./capo.md) — connect the community-driven capo SDK to RustFS with synchronous or asynchronous clients. +rclone](./rclone.md) — sync, mount, and serve RustFS buckets from the command line. +- [tusd](./tusd.md) — receive resumable uploads into a RustFS bucket over the tus protocol. +- [JuiceFS](./juicefs.md) — mount a POSIX filesystem backed by a RustFS bucket. +- [Nextcloud](./nextcloud.md) — use RustFS as S3 external storage for Nextcloud files. diff --git a/content/en/developer/integration/others/juicefs.md b/content/en/developer/integration/others/juicefs.md new file mode 100644 index 00000000..428b08ce --- /dev/null +++ b/content/en/developer/integration/others/juicefs.md @@ -0,0 +1,132 @@ +--- +title: "JuiceFS" +description: "Build a POSIX filesystem on RustFS with JuiceFS S3 object storage." +--- + +This guide connects [JuiceFS](https://github.com/juicedata/juicefs) — the cloud-native distributed POSIX filesystem — to **RustFS** as its object storage backend. You will format a volume whose data chunks live in a RustFS bucket, mount it locally, and read and write files through the mount. The workflow was verified with `juicefs v1.3.1` (community edition, SQLite metadata engine) against `rustfs/rustfs-x86-musl:v2.3.1`. + +You need the JuiceFS binary, a metadata engine, and FUSE (`fuse3` on Linux). This deployment is intended for local integration testing, not production. + +## Architecture + +```mermaid +flowchart LR + Mount["/mnt/jfs"] -->|"POSIX"| JuiceFS["JuiceFS client"] + JuiceFS -->|"metadata"| Meta["SQLite / Redis"] + JuiceFS -->|"data chunks"| RustFS["RustFS :9000"] +``` + +JuiceFS splits every file into chunks and stores them as objects under `chunks/` in the bucket, while the metadata engine tracks names, inodes, and layout. The filesystem behaves like a local disk but holds no data locally. + +## 1. Format the volume + +Create the bucket and format a JuiceFS volume backed by RustFS, replacing all connection placeholders. The bucket URL carries the endpoint, which selects path-style addressing: + +```bash +rc mb rustfs/jfs-demo + +juicefs format \ + --storage s3 \ + --bucket http://:9000/jfs-demo \ + --access-key \ + --secret-key \ + sqlite3:///opt/juicefs/jfs.db \ + rustfs-jfs +``` + +```text +Data use s3://:9000/jfs-demo/rustfs-jfs/ + OK, rustfs-jfs is ready +``` + +`sqlite3:///opt/juicefs/jfs.db` is the metadata engine for this test. In production, use Redis, MySQL, or PostgreSQL instead so multiple clients can mount the same volume. + +## 2. Mount the volume + +Mount the filesystem with the same metadata URL: + +```bash +mkdir -p /mnt/jfs +juicefs mount -d sqlite3:///opt/juicefs/jfs.db /mnt/jfs +``` + +```text +OK, rustfs-jfs is ready at /mnt/jfs +``` + +The `-d` flag runs the mount in the background. The volume is now a POSIX filesystem. + +## 3. Read and write files + +Use the mount like any other directory: + +```bash +echo "hello rustfs jfs" > /mnt/jfs/hello.txt +dd if=/dev/urandom of=/mnt/jfs/blob.bin bs=1M count=3 +mkdir -p /mnt/jfs/dir1 && echo nested > /mnt/jfs/dir1/nested.txt +cat /mnt/jfs/hello.txt +``` + +```text +hello rustfs jfs +``` + +Inspect the volume with `juicefs info`: + +```text +/mnt/jfs : + inode: 1 + files: 2 + dirs: 2 + length: 3.00 MiB +``` + +## 4. Verify chunks in RustFS + +List the bucket prefixes: + +```bash +rc ls rustfs/jfs-demo/ -r +``` + +Every file was split into content-addressed chunk objects: + +```text +rustfs-jfs/chunks/0/0/1_0_17 +rustfs-jfs/chunks/0/0/3_0_3145728 +rustfs-jfs/chunks/0/0/4_0_7 +``` + +The chunk name encodes the inode, chunk index, and size — for example `3_0_3145728` is the 3 MiB file written in step 3. + +![JuiceFS data chunks stored in the RustFS Console](./images/rustfs-juicefs-chunks.png) + +## 5. Stop or reset + +Unmount the volume, then optionally wipe the volume metadata and bucket data: + +```bash +juicefs umount /mnt/jfs +juicefs destroy --force sqlite3:///opt/juicefs/jfs.db rustfs-jfs +rc rm rustfs/jfs-demo/ --recursive --force +``` + +## Troubleshooting + +### `unknown option: --daemon` + +The background flag is a single dash: `juicefs mount -d`. Running without it keeps the mount in the foreground (useful for debugging). + +### `fusermount3: mount failed: Permission denied` + +Mounting requires the FUSE device. Inside a container, add `--device /dev/fuse --cap-add SYS_ADMIN` (or `--privileged`); on a host, install `fuse3` and confirm `/dev/fuse` exists. + +### Mount hangs or fails to reach storage + +The client must reach both the metadata engine and the bucket endpoint. Because the bucket URL embeds the endpoint, verify it from the mounting host with `curl` before formatting. + +## Next steps + +- Review [S3 compatibility notes](/administration/protocols/s3) before adopting additional JuiceFS storage backends. +- Create dedicated production credentials with [Access Key Management](/security-compliance/iam/access-token). +- Follow the [JuiceFS documentation](https://juicefs.com/docs/community/quick_start_guide/) to switch the metadata engine to Redis and mount the volume from multiple clients. diff --git a/content/en/developer/integration/others/meta.json b/content/en/developer/integration/others/meta.json index 65115f77..745fde1c 100644 --- a/content/en/developer/integration/others/meta.json +++ b/content/en/developer/integration/others/meta.json @@ -1,6 +1,10 @@ { "title": "Others", "pages": [ - "capo" + "capo", + "juicefs", + "nextcloud", + "rclone", + "tusd" ] } diff --git a/content/en/developer/integration/others/nextcloud.md b/content/en/developer/integration/others/nextcloud.md new file mode 100644 index 00000000..f506d75c --- /dev/null +++ b/content/en/developer/integration/others/nextcloud.md @@ -0,0 +1,145 @@ +--- +title: "Nextcloud" +description: "Use RustFS as S3 external storage for Nextcloud files." +--- + +This guide connects [Nextcloud](https://github.com/nextcloud/server) — the self-hosted content collaboration platform — to **RustFS** through its External Storage app with the S3 backend. You will enable `files_external`, mount a RustFS bucket into every user's files view, and upload a file through WebDAV that lands directly in the bucket. The workflow was verified with `nextcloud:32.0.15` (SQLite, single container) against `rustfs/rustfs-x86-musl:v2.3.1`. + +You need Docker, or an existing Nextcloud instance with `occ` access. This deployment is intended for local integration testing, not production. + +## Architecture + +```mermaid +flowchart LR + User["Browser / WebDAV"] --> Nextcloud["Nextcloud"] + Nextcloud -->|"files_external (S3)"| RustFS["RustFS :9000"] +``` + +Nextcloud proxies file operations on the mount point to the S3 backend. Objects are stored under their mount-relative paths, so the bucket mirrors the names users see. + +## 1. Install Nextcloud + +Run Nextcloud with an admin account, replacing all connection placeholders. SQLite keeps the test self-contained; use MariaDB or PostgreSQL in production: + +```bash +docker run -d --name nextcloud --network oo-rustfs_default -p 8080:80 \ + -e NEXTCLOUD_ADMIN_USER= \ + -e NEXTCLOUD_ADMIN_PASSWORD= \ + nextcloud:32.0.15 +``` + +If the web UI still shows the installer after startup, finish it manually: + +```bash +docker exec -u www-data nextcloud php occ maintenance:install \ + --admin-user --admin-password +``` + +## 2. Enable the External Storage app + +The `files_external` app ships with Nextcloud but starts disabled, and its `occ` commands only exist once the app is enabled: + +```bash +docker exec -u www-data nextcloud php occ app:enable files_external +``` + +```text +files_external 1.24.1 enabled +``` + +## 3. Mount the RustFS bucket + +Create an external storage of backend type `amazons3` with the `amazons3::accesskey` authentication backend. Replace all connection placeholders: + +```bash +docker exec -u www-data nextcloud php occ files_external:create \ + /rustfs amazons3 amazons3::accesskey \ + --user \ + --config bucket= \ + --config hostname= \ + --config port=9000 \ + --config use_ssl=false \ + --config use_path_style=true \ + --config key= \ + --config secret= +``` + +```text +Storage created with id 1 +``` + +The mount point `/rustfs` appears in the files view of the given user. `use_path_style=true` is required for a non-AWS endpoint. Check the connection before using it: + +```bash +docker exec -u www-data nextcloud php occ files_external:verify 1 +``` + +```text + - status: ok + - code: 0 +``` + +## 4. Upload a file and verify + +Upload through the WebDAV endpoint, which writes through the external storage: + +```bash +echo "nextcloud writes to rustfs" > /tmp/nc-demo.txt + +curl -u : \ + -T /tmp/nc-demo.txt \ + http://localhost:8080/remote.php/dav/files//rustfs/nc-demo.txt \ + -o /dev/null -w "%{http_code}\n" +``` + +```text +201 +``` + +Read it back through the same path, then confirm the object in RustFS: + +```bash +rc ls rustfs// -r +``` + +```text +[2026-09-29 13:42:48] 27 B nc-demo.txt +``` + +The object key equals the path inside the mount, so files uploaded through Nextcloud can also be read directly with any S3 client. + +![Nextcloud file stored in the RustFS Console](./images/rustfs-nextcloud-file.png) + +## 5. Stop or reset + +To remove the mount without touching the bucket: + +```bash +docker exec -u www-data nextcloud php occ files_external:delete 1 +``` + +To delete the bucket contents: + +```bash +rc rm rustfs// --recursive --force +``` + +## Troubleshooting + +### `There are no commands defined in the "files_external" namespace` + +The app is not enabled yet. Run `occ app:enable files_external` first; the `occ files_external:*` commands only register afterwards. + +### `Not enough arguments (missing: "authentication_backend")` + +`files_external:create` takes the storage backend and the authentication backend as two separate arguments: `amazons3 amazons3::accesskey`. The backend identifiers are listed by `occ files_external:backends`. + +### Mount shows but is empty, or uploads fail + +Confirm `hostname` is reachable from the Nextcloud container (use the container network name, not `localhost`), `use_path_style` is `true`, and the bucket exists. `occ files_external:verify ` reports the exact connection error. + +## Next steps + +- Review [S3 compatibility notes](/administration/protocols/s3) before adopting additional external storage backends. +- Create dedicated production credentials with [Access Key Management](/security-compliance/iam/access-token). +- Follow the [Nextcloud external storage documentation](https://docs.nextcloud.com/server/latest/admin_manual/configuration_files/external_storage_configuration_gui.html) to share the mount with groups and enable versioning. diff --git a/content/en/developer/integration/others/rclone.md b/content/en/developer/integration/others/rclone.md new file mode 100644 index 00000000..34f9fe4a --- /dev/null +++ b/content/en/developer/integration/others/rclone.md @@ -0,0 +1,144 @@ +--- +title: "rclone" +description: "Sync, mount, and serve RustFS buckets with rclone over its S3-compatible API." +--- + +This guide connects [rclone](https://github.com/rclone/rclone) — the command-line tool for syncing files to and from cloud storage — to **RustFS** through its S3 backend. You will configure an S3 remote for RustFS, copy and sync files, read objects back, publish a bucket over HTTP with `rclone serve`, and mount the bucket as a local filesystem with `rclone mount`. The workflow was verified with `rclone v1.75.1` against `rustfs/rustfs-x86-musl:v2.3.1`. + +You need Docker, or a local rclone binary. This deployment is intended for local integration testing, not production. + +## Architecture + +```mermaid +flowchart LR + Files["Local files"] -->|"copy / sync"| Remote["rclone S3 remote"] + Remote -->|"S3 API"| RustFS["RustFS :9000"] + RustFS -->|"mount / serve"| Client["FUSE mount / HTTP clients"] +``` + +One remote definition drives every rclone command: data transfer, mounting, and serving all use the same S3 connection. + +## 1. Configure the remote + +Create an rclone config file with an S3 remote for RustFS, replacing all connection placeholders. The `Other` provider disables AWS-specific behavior, and path-style addressing is used automatically for custom endpoints: + +```ini title="rclone.conf" +[rustfs] +type = s3 +provider = Other +access_key_id = +secret_access_key = +endpoint = http://:9000 +region = us-east-1 +``` + +## 2. Copy and read objects + +Create the bucket and upload a directory with `rclone copy`: + +```bash +rc mb rustfs/rclone-demo +rclone copy /data rustfs:rclone-demo/seed +``` + +List and read back: + +```bash +rclone ls rustfs:rclone-demo/seed +rclone cat rustfs:rclone-demo/seed/hello.txt +``` + +```text + 3145728 blob.bin + 18 hello.txt +hello from rclone +``` + +`rclone lsd rustfs:` lists every bucket on the endpoint. + +## 3. Sync a directory + +`rclone sync` makes the destination identical to the source, including deletions. Remove a local file and sync: + +```bash +rm /data/hello.txt +rclone sync /data rustfs:rclone-demo/seed +rclone lsf rustfs:rclone-demo/seed +``` + +```text +blob.bin +``` + +`hello.txt` disappears from the bucket. Add `--dry-run` first to preview the changes without touching the bucket. + +## 4. Serve a bucket over HTTP + +Publish the bucket contents as an HTTP file server: + +```bash +rclone serve http --addr 0.0.0.0:8080 rustfs:rclone-demo/seed +``` + +Any HTTP client can now download objects: + +```bash +curl -s http://localhost:8080/blob.bin -o /dev/null -w "%{http_code} %{size_download} bytes\n" +``` + +```text +200 3145728 bytes +``` + +`rclone serve` also supports WebDAV, SFTP, and S3 endpoints over the same remote. + +## 5. Mount the bucket as a filesystem + +With FUSE available, mount the bucket locally and use it like a directory: + +```bash +rclone mount rustfs:rclone-demo /mnt/rclone --daemon +ls /mnt/rclone/seed +echo test > /mnt/rclone/write-test.txt +cat /mnt/rclone/write-test.txt +``` + +Files written through the mount appear in RustFS as regular objects: + +```bash +rc ls rustfs/rclone-demo/ -r +``` + +```text +[2026-09-29 11:12:41] 5 B write-test.txt +``` + +Unmount with `fusermount -u /mnt/rclone` when finished. + +## 6. Stop or reset + +rclone holds no server-side state. To delete the demo data: + +```bash +rclone purge rustfs:rclone-demo +``` + +## Troubleshooting + +### `Access Denied` or empty listings + +Confirm `endpoint` includes the scheme and that the key pair matches a RustFS access key. The `region` value is required by the S3 signer even though RustFS ignores it; keep `us-east-1`. + +### Mount fails with `fusermount3: mount failed: Permission denied` + +Mounting needs the FUSE device and elevated privileges. Inside a container, run with `--device /dev/fuse --cap-add SYS_ADMIN` — and use `--privileged` if the mount helper still fails. On a host, verify that `fuse3` is installed and `/dev/fuse` exists. + +### Sync deleted nothing on the destination + +`rclone copy` never deletes. Only `rclone sync` (or `rclone delete`) removes destination objects, and `--dry-run` is the safe way to preview either. + +## Next steps + +- Review [S3 compatibility notes](/administration/protocols/s3) before adopting additional rclone backends. +- Create dedicated production credentials with [Access Key Management](/security-compliance/iam/access-token). +- Follow the [rclone S3 documentation](https://rclone.org/s3/) for flags such as `--transfers`, bandwidth limits, and crypt overlays. diff --git a/content/en/developer/integration/others/tusd.md b/content/en/developer/integration/others/tusd.md new file mode 100644 index 00000000..99f3c588 --- /dev/null +++ b/content/en/developer/integration/others/tusd.md @@ -0,0 +1,170 @@ +--- +title: "tusd" +description: "Receive resumable uploads into RustFS with the tusd server's S3 backend." +--- + +This guide connects [tusd](https://github.com/tus/tusd) — the official reference implementation of the tus resumable-upload protocol — to **RustFS** as its S3 storage backend. You will run tusd against a RustFS bucket, create an upload with the tus protocol, send the file in two chunks with an interruption in between, resume from the reported offset, and verify the assembled object in the bucket. The workflow was verified with `tusproject/tusd:v2.10.1` against `rustfs/rustfs-x86-musl:v2.3.1`. + +You need Docker, or a local tusd binary. This deployment is intended for local integration testing, not production. + +## Architecture + +```mermaid +flowchart LR + Client["tus client"] -->|"POST / PATCH / HEAD"| tusd["tusd :8080"] + tusd -->|"multipart upload"| RustFS["RustFS :9000"] +``` + +tusd stores each in-progress upload as S3 multipart parts in the bucket. A client that loses its connection asks the server for the last committed offset with `HEAD` and continues from there — the data already received is never sent twice. + +## 1. Run tusd + +Create the bucket and start tusd with the S3 backend, replacing all connection placeholders. The AWS region must be set even though RustFS ignores it: + +```bash +rc mb rustfs/tus-uploads + +docker run -d --name tusd --network oo-rustfs_default -p 8080:8080 \ + -e AWS_ACCESS_KEY_ID= \ + -e AWS_SECRET_ACCESS_KEY= \ + -e AWS_REGION=us-east-1 \ + tusproject/tusd:latest \ + -s3-bucket tus-uploads \ + -s3-endpoint http://:9000 +``` + +Check that the server is healthy: + +```bash +curl -s -o /dev/null -w "%{http_code}\n" http://localhost:8080/health +``` + +```text +200 +``` + +## 2. Create the upload + +Create a 6 MiB upload and read the `Location` header: + +```bash +curl -s -D - -o /dev/null -X POST http://localhost:8080/files/ \ + -H "Upload-Length: 6291456" -H "Tus-Resumable: 1.0.0" \ + | grep -i "^Location:" +``` + +```text +Location: http://localhost:8080/files/b7338250daa9...+NGVmNDRhZjEt... +``` + +The upload URL contains the file ID and a message-authentication tag. Strip the scheme and host before re-sending it (the server echoes whatever `Host` it received, which may not be reachable from your next client). + +## 3. Upload in chunks with an interruption + +Send the first 2.5 MB, then stop — this is the point where a mobile client would lose its connection: + +```bash +head -c 2500000 demo.bin > part1.bin + +curl -s -o /dev/null -w "%{http_code}\n" -X PATCH "http://localhost:8080${LOC}" \ + -H "Upload-Offset: 0" -H "Tus-Resumable: 1.0.0" \ + -H "Content-Type: application/offset+octet-stream" \ + --data-binary @part1.bin +``` + +```text +204 +``` + +Ask the server how much it actually has — this is the resumable-upload core: + +```bash +curl -s -X HEAD "http://localhost:8080${LOC}" \ + -H "Tus-Resumable: 1.0.0" -D - -o /dev/null | grep -i upload-offset +``` + +```text +Upload-Offset: 2500000 +``` + +## 4. Resume and finish + +Continue from offset 2500000 with the remaining bytes: + +```bash +tail -c 3791456 demo.bin > part2.bin + +curl -s -o /dev/null -w "%{http_code}\n" -X PATCH "http://localhost:8080${LOC}" \ + -H "Upload-Offset: 2500000" -H "Tus-Resumable: 1.0.0" \ + -H "Content-Type: application/offset+octet-stream" \ + --data-binary @part2.bin +``` + +```text +204 +``` + +Download the finished upload through tusd and compare checksums with the source: + +```bash +curl -s -o download.bin "http://localhost:8080${LOC}" +sha1sum demo.bin download.bin +``` + +```text +d9016032ced6c7515b67a0c556e006c4b25a5858 demo.bin +d9016032ced6c7515b67a0c556e006c4b25a5858 download.bin +``` + +## 5. Verify objects in RustFS + +List the bucket: + +```bash +rc ls rustfs/tus-uploads/ -r +``` + +The bucket holds the assembled object plus one `.info` metadata file per upload — both live and finished: + +```text +b7338250daa9a1a79c1343502b57b28f 6 MiB +b7338250daa9a1a79c1343502b57b28f.info 378 B +``` + +The object key is the upload ID, and the object body is the uploaded file byte-for-byte — so any S3 client can read completed uploads directly from the bucket. + +![tus uploads stored in the RustFS Console](./images/rustfs-tus-uploads.png) + +## 6. Stop or reset + +To tear down the server while keeping the bucket objects: + +```bash +docker rm -f tusd +``` + +To delete the stored uploads: + +```bash +rc rm rustfs/tus-uploads/ --recursive --force +``` + +## Troubleshooting + +### `CreateMultipartUpload ... A region must be set when sending requests to S3` + +tusd builds its S3 client from the AWS environment, and the region is mandatory for endpoint resolution. Export `AWS_REGION=us-east-1` next to the credentials, as in step 1. + +### `PATCH` returns `404` or connects to the wrong host + +The `Location` URL echoes the `Host` header of the creation request. When your client and the server use different hostnames (container name versus published port), strip the scheme and host from the URL and send the path to the address the client can reach. + +### Upload disappears after server restart + +The S3 backend keeps `.info` files in the bucket, so uploads survive restarts. If you run tusd against an empty bucket that another process prunes, the metadata is lost — protect the `tus-uploads` prefix from cleanup jobs. + +## Next steps + +- Review [S3 compatibility notes](/administration/protocols/s3) before adopting additional tusd backends. +- Create dedicated production credentials with [Access Key Management](/security-compliance/iam/access-token). +- Follow the [tus protocol documentation](https://tus.io/protocols/resumable-upload) for creation-with-upload, termination, and checksum extensions that tusd supports on top of the core protocol. diff --git a/content/en/developer/integration/reverse-proxy/envoy.md b/content/en/developer/integration/reverse-proxy/envoy.md new file mode 100644 index 00000000..dc635a48 --- /dev/null +++ b/content/en/developer/integration/reverse-proxy/envoy.md @@ -0,0 +1,192 @@ +--- +title: "Envoy" +description: "Deploy RustFS behind Envoy with TLS-terminated routes for the S3 API and Console." +--- + +Use **Envoy** to terminate TLS and route separate hostnames to the RustFS S3 API and Console. This deployment runs Envoy and a single-node RustFS instance on one Docker network. You need Docker Engine, two DNS records, and a TLS certificate that covers both hostnames (the guide uses a self-signed certificate for testing). + +This guide uses these example hostnames: + +- `s3.example.com` for the S3 API +- `console.example.com` for the Console + +Replace them with hostnames that resolve to the Docker host. + +:::warning[Serve S3 from the root path] + +Do not publish the S3 API under a path such as `/s3/`. AWS Signature Version 4 includes the request path and host, so rewriting either value can invalidate signed requests. Envoy forwards the incoming `Host` header unchanged, which keeps signatures valid. + +::: + +## 1. Create the deployment directories + +Create directories for the Envoy configuration and TLS certificate: + +```bash +mkdir -p rustfs-envoy/certs +cd rustfs-envoy +``` + +For local testing, generate a self-signed certificate covering both hostnames: + +```bash +openssl req -x509 -newkey rsa:2048 -nodes \ + -keyout certs/privkey.pem -out certs/fullchain.pem -days 30 \ + -subj "/CN=*.example.com" \ + -addext "subjectAltName=DNS:s3.example.com,DNS:console.example.com" +chmod 644 certs/privkey.pem +``` + +Make the certificate readable by the non-root user the Envoy image runs as — a `600` private key produces a misleading `Failed to load incomplete private key` error. + +## 2. Configure Envoy + +Create the configuration with an HTTPS listener and two virtual hosts. Route timeouts are disabled (`timeout: 0s`) so long streaming S3 uploads are not cut off: + +```yaml title="envoy.yaml" +static_resources: + listeners: + - name: https + address: {socket_address: {address: 0.0.0.0, port_value: 8443}} + filter_chains: + - transport_socket: + name: envoy.transport_sockets.tls + typed_config: + "@type": type.googleapis.com/envoy.extensions.transport_sockets.tls.v3.DownstreamTlsContext + common_tls_context: + tls_certificates: + - certificate_chain: {filename: /certs/fullchain.pem} + private_key: {filename: /certs/privkey.pem} + filters: + - name: envoy.filters.network.http_connection_manager + typed_config: + "@type": type.googleapis.com/envoy.extensions.filters.network.http_connection_manager.v3.HttpConnectionManager + stat_prefix: rustfs_https + route_config: + virtual_hosts: + - name: s3 + domains: ["s3.example.com", "s3.example.com:*"] + routes: + - match: {prefix: "/"} + route: {cluster: rustfs_s3, timeout: 0s} + - name: console + domains: ["console.example.com", "console.example.com:*"] + routes: + - match: {prefix: "/"} + route: {cluster: rustfs_console, timeout: 0s} + http_filters: + - name: envoy.filters.http.router + typed_config: + "@type": type.googleapis.com/envoy.extensions.filters.http.router.v3.Router + clusters: + - name: rustfs_s3 + connect_timeout: 5s + type: STRICT_DNS + lb_policy: ROUND_ROBIN + load_assignment: + cluster_name: rustfs_s3 + endpoints: + - lb_endpoints: + - endpoint: {address: {socket_address: {address: rustfs, port_value: 9000}}} + - name: rustfs_console + connect_timeout: 5s + type: STRICT_DNS + lb_policy: ROUND_ROBIN + load_assignment: + cluster_name: rustfs_console + endpoints: + - lb_endpoints: + - endpoint: {address: {socket_address: {address: rustfs, port_value: 9001}}} +``` + +The `host:*` domain entries matter: clients send `Host: s3.example.com:8443` on non-standard ports, and Envoy matches the authority including the port. + +## 3. Start Envoy + +Run Envoy on the same Docker network as RustFS, publishing only the proxy port: + +```bash +docker run -d --name envoy --network oo-rustfs_default -p 8443:8443 \ + -v "$PWD/envoy.yaml":/envoy.yaml:ro \ + -v "$PWD/certs":/certs:ro \ + envoyproxy/envoy:v1.34-latest -c /envoy.yaml +``` + +## 4. Verify both endpoints + +Point the example hostnames at the proxy with `curl --resolve` (in production, DNS does this): + +```bash +curl -sk --resolve s3.example.com:8443:127.0.0.1 \ + https://s3.example.com:8443/health/ready -o /dev/null -w "s3 api: %{http_code}\n" + +curl -sk --resolve console.example.com:8443:127.0.0.1 \ + https://console.example.com:8443/rustfs/console/ -o /dev/null -w "console: %{http_code}\n" +``` + +```text +s3 api: 200 +console: 200 +``` + +`-k` skips certificate validation because the certificate is self-signed; with a trusted certificate, drop it. + +## 5. Send signed S3 requests through Envoy + +Point any S3 client at the proxy as if it were RustFS. Configure the client with `https://s3.example.com:8443` as the endpoint and path-style addressing; when the proxy certificate is trusted, signed AWS Signature Version 4 requests pass through unchanged. For a quick test over plain HTTP, add an HTTP listener on port 8080 with the same virtual-host routing as the HTTPS listener, then use the endpoint `http://s3.example.com:8080`: + +```bash +rc alias set rustfs-envoy http://s3.example.com:8080 +rc ls rustfs-envoy/rclone-demo/ +``` + +```text +[ ] 0B seed/ +[2026-09-29 11:12:41] 5 B write-test.txt +``` + +The signatures validate because Envoy forwards the original `Host` header to RustFS. + +## Multi-node backends + +For a distributed RustFS deployment, add every node to the S3 cluster: + +```yaml title="envoy.yaml" + - name: rustfs_s3 + connect_timeout: 5s + type: STRICT_DNS + lb_policy: ROUND_ROBIN + load_assignment: + cluster_name: rustfs_s3 + endpoints: + - lb_endpoints: + - endpoint: {address: {socket_address: {address: node1, port_value: 9000}}} + - lb_endpoints: + - endpoint: {address: {socket_address: {address: node2, port_value: 9000}}} + - lb_endpoints: + - endpoint: {address: {socket_address: {address: node3, port_value: 9000}}} + - lb_endpoints: + - endpoint: {address: {socket_address: {address: node4, port_value: 9000}}} +``` + +The Console cluster follows the same pattern on port `9001`. + +## Troubleshooting + +### `Failed to load incomplete private key from path` + +The Envoy container runs as a non-root user and cannot read a `600` root-owned key. `chmod 644` the key files (or chown them to the container user, UID `1001` in the official image). + +### Routes return `404` with the correct hostnames + +Envoy matches the authority including the port. Add the `host:*` variants to each virtual host's `domains` list, as in the configuration above. + +### `Access Denied` from RustFS on proxied requests + +Confirm the proxy is not rewriting the path or the `Host` header. Signed requests must reach RustFS with the host the client signed for. + +## Next steps + +- [Configure an S3 client](/developer/examples/aws-cli) +- [Enable virtual-hosted-style bucket URLs](/integration/virtual) +- [Review health and readiness endpoints](/operations/status-check) diff --git a/content/en/developer/integration/reverse-proxy/index.md b/content/en/developer/integration/reverse-proxy/index.md index eff7324d..397cdad6 100644 --- a/content/en/developer/integration/reverse-proxy/index.md +++ b/content/en/developer/integration/reverse-proxy/index.md @@ -13,6 +13,7 @@ We recommend using separate hostnames for the S3 API on port `9000` and the Cons - [Traefik](./traefik.md) - [Caddy](./caddy.md) - [HAProxy](./haproxy.md) +- [Envoy](./envoy.md) - [Apache HTTP Server](./httpd.md) ## Related configuration diff --git a/content/en/developer/integration/reverse-proxy/meta.json b/content/en/developer/integration/reverse-proxy/meta.json index adeb173e..47c8243f 100644 --- a/content/en/developer/integration/reverse-proxy/meta.json +++ b/content/en/developer/integration/reverse-proxy/meta.json @@ -4,7 +4,8 @@ "nginx", "traefik", "caddy", + "envoy", "haproxy", "httpd" ] -} \ No newline at end of file +} diff --git a/content/fr/developer/integration/ai/images/rustfs-vllm-models.png b/content/fr/developer/integration/ai/images/rustfs-vllm-models.png new file mode 100644 index 00000000..baa18dc1 Binary files /dev/null and b/content/fr/developer/integration/ai/images/rustfs-vllm-models.png differ diff --git a/content/fr/developer/integration/ai/index.md b/content/fr/developer/integration/ai/index.md index 37fa7892..7837dbaa 100644 --- a/content/fr/developer/integration/ai/index.md +++ b/content/fr/developer/integration/ai/index.md @@ -8,5 +8,6 @@ Utilisez **RustFS** comme couche de stockage objet pour les plateformes d'IA et ## Plateformes - [Ray](./ray.md) +- [vLLM](./vllm.md) Conservez les jeux de données et les checkpoints dans des buckets dédiés et limitez les identifiants aux opérations de bucket requises. diff --git a/content/fr/developer/integration/ai/meta.json b/content/fr/developer/integration/ai/meta.json index 876f2b47..639b7f23 100644 --- a/content/fr/developer/integration/ai/meta.json +++ b/content/fr/developer/integration/ai/meta.json @@ -1,6 +1,7 @@ { "title": "IA", "pages": [ - "ray" + "ray", + "vllm" ] } diff --git a/content/fr/developer/integration/ai/vllm.md b/content/fr/developer/integration/ai/vllm.md new file mode 100644 index 00000000..c471096c --- /dev/null +++ b/content/fr/developer/integration/ai/vllm.md @@ -0,0 +1,154 @@ +--- +title: "vLLM" +description: "Serve LLM inference with vLLM loading model weights stored in RustFS." +--- + +This guide connects [vLLM](https://github.com/vllm-project/vllm) — the high-throughput LLM inference engine — to **RustFS** as its model-weight store. You will upload a model into a RustFS bucket, expose the bucket to the vLLM host through an rclone mount, and serve the model with the OpenAI-compatible API. The workflow was verified with `vllm/vllm-openai-cpu` (vLLM 0.30.0) serving `facebook/opt-125m` from a RustFS bucket backed by `rustfs/rustfs-x86-musl:v2.3.1`, on a CPU-only host. + +You need Docker and an rclone binary on the host that runs vLLM. This deployment is intended for local integration testing, not production. + +## Architecture + +```mermaid +flowchart LR + Upload["rclone copy"] -->|"weights"| RustFS["RustFS :9000"] + RustFS -->|"rclone mount"| Mount["/mnt/vllm-models"] + Mount -->|"weight load"| vLLM["vLLM :8000"] + Client["OpenAI SDK / curl"] -->|"completions"| vLLM +``` + +The bucket is the single copy of the model. Hosts that serve the model mount the bucket read-only, so every node pulls weights from RustFS and no local model store exists to drift. + +:::note[Why an rclone mount] + +vLLM 0.30 loads `s3://` model paths through the RunAI model streamer, whose ranged reads currently fail against custom S3 endpoints such as RustFS (the loader errors with `File access error` on any non-zero offset). Mounting the bucket as a filesystem is the verified way to keep the weights in RustFS while vLLM reads them as local files. + +::: + +## 1. Upload the model to RustFS + +Create the bucket and copy model weights into it, replacing all connection placeholders: + +```ini title="rclone.conf" +[rustfs] +type = s3 +provider = Other +access_key_id = +secret_access_key = +endpoint = http://:9000 +region = us-east-1 +``` + +```bash +rc mb rustfs/vllm-models +rclone copy ./opt-125m rustfs:vllm-models/opt-125m --transfers 4 +``` + +Any Hugging Face layout works — `config.json`, the tokenizer files, and the weight files (`model.safetensors` or `pytorch_model.bin`). Keep one model per prefix so several models can share the bucket. + +## 2. Mount the bucket on the vLLM host + +On the machine that runs vLLM, mount the bucket read-only for clients with `--allow-other`: + +```bash +mkdir -p /mnt/vllm-models +rclone mount rustfs:vllm-models /mnt/vllm-models \ + --allow-other --daemon +ls /mnt/vllm-models/opt-125m/ +``` + +```text +config.json merges.txt model.safetensors tokenizer.json vocab.json +``` + +## 3. Run vLLM + +Start the CPU image against the mounted weights: + +```bash +docker run -d --name vllm -p 8000:8000 --shm-size=2g \ + -v /mnt/vllm-models:/models:ro \ + vllm/vllm-openai-cpu:latest \ + --model /models/opt-125m --served-model-name opt-125m \ + --dtype float32 --max-model-len 256 --gpu-memory-utilization 0.15 +``` + +vLLM reads the weights through the mount — the container stays stateless and the model lives in RustFS. Wait for the server to come up: + +```bash +curl -s http://localhost:8000/v1/models | head -c 200 +``` + +```text +{"object":"list","data":[{"id":"opt-125m","object":"model","created":...,"root":"/models/opt-125m",...}]} +``` + +`--gpu-memory-utilization` controls the fraction of RAM reserved for the KV cache on the CPU backend; lower it on small hosts. `--dtype float32` matches what the CPU attention kernels support for this model. + +## 4. Run inference + +Send an OpenAI-compatible completion request: + +```bash +curl -s http://localhost:8000/v1/completions \ + -H "Content-Type: application/json" \ + -d '{"model": "opt-125m", "prompt": "RustFS is", "max_tokens": 12, "temperature": 0}' +``` + +```json +{"id":"cmpl-...","object":"text_completion","model":"opt-125m", + "choices":[{"index":0,"text":" a great tool for building your own server. It's a", + "finish_reason":"length",...}]} +``` + +The request is standard OpenAI schema, so the Python client works unchanged: + +```python +from openai import OpenAI + +client = OpenAI(base_url="http://localhost:8000/v1", api_key="EMPTY") +print(client.completions.create( + model="opt-125m", prompt="RustFS is", max_tokens=12, temperature=0, +).choices[0].text) +``` + +![vLLM model weights stored in the RustFS Console](./images/rustfs-vllm-models.png) + +## 5. Stop or reset + +To tear down the demo while keeping the bucket objects: + +```bash +docker rm -f vllm +fusermount -u /mnt/vllm-models +``` + +To delete the stored model: + +```bash +rclone purge rustfs:vllm-models +``` + +## Troubleshooting + +### `Cannot find any model weights with /models/...` + +The mount had a stale directory cache or the weight files never made it to the bucket. Re-run `rclone copy` and confirm the files through the mount with `ls` before starting vLLM. A short `--dir-cache-time` (for example `10s`) helps while you iterate. + +### `Unsupported CPU attention configuration: head_dim=...` + +vLLM's CPU kernels support a fixed set of head dimensions. Tiny test models such as `hf-internal-testing/tiny-random-*` use exotic shapes that fail at request time — use a real small model such as `facebook/opt-125m`. + +### `Insufficient space in /dev/shm` + +vLLM's CPU engine exchanges tensors through shared memory. Run the container with `--shm-size=2g` (or `--ipc=host`). + +### Server exits with `Available memory on node 0 ... is less than desired CPU memory utilization` + +The default KV-cache reservation is 90% of system RAM. Lower it with `--gpu-memory-utilization 0.15` (the flag applies to the CPU backend as a memory fraction despite its name). + +## Next steps + +- Review [S3 compatibility notes](/administration/protocols/s3) before adopting additional serving setups. +- Create dedicated production credentials with [Access Key Management](/security-compliance/iam/access-token). +- Follow the [vLLM documentation](https://docs.vllm.ai/en/latest/serving/openai_compatible_server.html) for chat templates, tensor parallelism, and quantized weights on top of the same bucket-backed model store. diff --git a/content/fr/developer/integration/big-data/airflow.md b/content/fr/developer/integration/big-data/airflow.md new file mode 100644 index 00000000..7f55e219 --- /dev/null +++ b/content/fr/developer/integration/big-data/airflow.md @@ -0,0 +1,170 @@ +--- +title: "Airflow" +description: "Move data between Airflow DAGs and RustFS with the Amazon S3 provider." +--- + +This guide connects [Apache Airflow](https://github.com/apache/airflow) — the workflow orchestration platform — to **RustFS** through the Amazon S3 provider's hooks, operators, and sensors. You will register a custom-endpoint connection, run a DAG that writes an object to a RustFS bucket, waits for a key with `S3KeySensor`, and reads the object back with `S3Hook`. The workflow was verified with `apache/airflow:3.3.2` (standalone, SequentialExecutor) and the `apache-airflow-providers-amazon` provider against `rustfs/rustfs-x86-musl:v2.3.1`. + +You need Docker, or an existing Airflow installation. This deployment is intended for local integration testing, not production. + +## Architecture + +```mermaid +flowchart LR + Scheduler["Airflow scheduler"] -->|"tasks"| Hook["S3Hook / operators"] + Hook -->|"S3 API"| RustFS["RustFS :9000"] + Sensor["S3KeySensor"] -->|"poll key"| RustFS +``` + +Every S3 interaction inside a DAG goes through the provider's S3 client, pointed at RustFS by the connection's `endpoint_url`. Operators, sensors, and hooks share the same connection object. + +## 1. Run Airflow + +Start a standalone instance with examples disabled, and create the demo bucket: + +```bash +docker run -d --name airflow --network oo-rustfs_default -p 8080:8080 \ + -e AIRFLOW__CORE__LOAD_EXAMPLES=False \ + -v "$PWD/dags":/opt/airflow/dags \ + apache/airflow:3.3.2 standalone + +rc mb rustfs/airflow-demo +``` + +The image ships with all providers preinstalled, including `apache-airflow-providers-amazon`. + +## 2. Register the RustFS connection + +The S3 provider reads its endpoint from the connection's extra field. Replace all connection placeholders: + +```bash +docker exec airflow airflow connections add rustfs \ + --conn-type aws \ + --conn-extra '{"endpoint_url": "http://:9000", "region_name": "us-east-1", "aws_access_key_id": "", "aws_secret_access_key": ""}' +``` + +The keys `aws_access_key_id` and `aws_secret_access_key` inside `--conn-extra` supply credentials; `endpoint_url` redirects the boto3 client from AWS to RustFS. + +## 3. Write the DAG + +The DAG writes an object with an operator, waits for the key with a sensor, and reads it back with the hook: + +```python title="rustfs_demo.py" +import datetime + +from airflow.providers.amazon.aws.hooks.s3 import S3Hook +from airflow.providers.amazon.aws.operators.s3 import S3CreateObjectOperator +from airflow.providers.amazon.aws.sensors.s3 import S3KeySensor +from airflow.sdk import dag, task + +@dag( + schedule=None, + start_date=datetime.datetime(2026, 1, 1), + catchup=False, + tags=["rustfs"], +) +def rustfs_demo(): + create = S3CreateObjectOperator( + task_id="write_object", + s3_bucket="airflow-demo", + s3_key="dags/airflow-put.txt", + data="written by airflow to rustfs", + aws_conn_id="rustfs", + replace=True, + ) + + wait = S3KeySensor( + task_id="wait_for_object", + bucket_key="dags/airflow-put.txt", + bucket_name="airflow-demo", + aws_conn_id="rustfs", + timeout=120, + poke_interval=10, + mode="reschedule", + ) + + @task + def read_object(): + hook = S3Hook(aws_conn_id="rustfs") + body = hook.read_key(key="dags/airflow-put.txt", bucket_name="airflow-demo") + print("read back:", body) + assert body == "written by airflow to rustfs" + + create >> [wait, read_object()] + +rustfs_demo() +``` + +Note the import paths: `S3CreateObjectOperator` lives in the `operators` module while `S3KeySensor` lives in the `sensors` module — importing both from one place fails. + +## 4. Unpause and trigger + +New DAGs start paused, and a trigger fired while paused stays queued forever. Unpause first, then trigger: + +```bash +docker exec airflow airflow dags unpause rustfs_demo +docker exec airflow airflow dags trigger rustfs_demo +``` + +Watch the run finish: + +```bash +docker exec airflow airflow dags list-runs rustfs_demo | head -3 +``` + +```text +dag_id run_id state +rustfs_demo manual__2026-09-29T13:15:57.332688+00:00 success +``` + +All three tasks succeed: `write_object`, `wait_for_object`, and `read_object`. + +## 5. Verify objects in RustFS + +List the bucket prefix: + +```bash +rc ls rustfs/airflow-demo/ -r +rc cat rustfs/airflow-demo/dags/airflow-put.txt +``` + +```text +[2026-09-29 13:16:01] 28 B dags/airflow-put.txt +written by airflow to rustfs +``` + +![Airflow object stored in the RustFS Console](./images/rustfs-airflow-object.png) + +## 6. Stop or reset + +To tear down the demo while keeping the bucket objects: + +```bash +docker rm -f airflow +``` + +To delete the stored data: + +```bash +rc rm rustfs/airflow-demo/ --recursive --force +``` + +## Troubleshooting + +### Dag runs stay `queued` after triggering + +The DAG is paused. New DAGs are paused by default in Airflow 3, and runs triggered in that state never execute. Run `airflow dags unpause rustfs_demo`; queued runs then start on their own. + +### `cannot import name 'S3KeySensor' from 'airflow.providers.amazon.aws.operators.s3'` + +The sensor lives in a separate module: `from airflow.providers.amazon.aws.sensors.s3 import S3KeySensor`. + +### Tasks fail with connection errors + +The `endpoint_url` must be reachable from the Airflow container — use the Docker network hostname for RustFS, not `localhost`. Airflow 3 serves its health endpoint under `/api/v2/monitor/health` if you need to check component status. + +## Next steps + +- Review [S3 compatibility notes](/administration/protocols/s3) before adopting additional provider hooks. +- Create dedicated production credentials with [Access Key Management](/security-compliance/iam/access-token). +- Follow the [Amazon provider documentation](https://airflow.apache.org/docs/apache-airflow-providers-amazon/stable/index.html) for transfer operators such as `S3ToLocalFilesystemOperator` and `LocalFilesystemToS3Operator`. diff --git a/content/fr/developer/integration/big-data/delta-lake.md b/content/fr/developer/integration/big-data/delta-lake.md new file mode 100644 index 00000000..b1bfae56 --- /dev/null +++ b/content/fr/developer/integration/big-data/delta-lake.md @@ -0,0 +1,125 @@ +--- +title: "Delta Lake" +description: "Write and read Delta tables on RustFS with delta-rs." +--- + +This guide connects [Delta Lake](https://github.com/delta-io/delta) — the open-source lakehouse table format — to **RustFS** through delta-rs, the Rust-native Delta implementation. You will write a Delta table to a RustFS bucket from Python, read it back with ACID transaction history, and confirm the `_delta_log` and Parquet files in the bucket. The workflow was verified with the `deltalake` Python package (delta-rs) and pandas against `rustfs/rustfs-x86-musl:v2.3.1`. + +You need Python 3.9 or newer. This deployment is intended for local integration testing, not production. + +## Architecture + +```mermaid +flowchart LR + DF["pandas DataFrame"] -->|"write_deltalake"| deltaRS["delta-rs"] + deltaRS -->|"Parquet + _delta_log"| RustFS["RustFS :9000"] + Query["DeltaTable"] -->|"read / time travel"| RustFS +``` + +delta-rs stores each table as Parquet files plus a transaction log (`_delta_log/`). All I/O goes through the `object_store` crate, configured with the same AWS environment variables as other S3 clients. + +## 1. Install the client + +```bash +pip install deltalake pandas pyarrow +``` + +`pyarrow` is required to convert pandas frames into Delta-compatible record batches. + +## 2. Write a Delta table + +Create the bucket and write a table, replacing all connection placeholders. `AWS_S3_ALLOW_UNSAFE_RENAME` is needed because RustFS does not provide copy-if-not-exists, which delta-rs otherwise uses for commit conflicts: + +```python title="delta_s3.py" +import pandas as pd +from deltalake import DeltaTable, write_deltalake + +storage_options = { + "AWS_ENDPOINT_URL": "http://:9000", + "AWS_ACCESS_KEY_ID": "", + "AWS_SECRET_ACCESS_KEY": "", + "AWS_REGION": "us-east-1", + "AWS_ALLOW_HTTP": "true", + "AWS_S3_ALLOW_UNSAFE_RENAME": "true", +} + +table = "s3:///events" +df = pd.DataFrame({"id": [1, 2, 3], "name": ["alpha", "beta", "gamma"]}) +write_deltalake(table, df, storage_options=storage_options) +print("written:", df.shape[0], "rows") +``` + +```text +written: 3 rows +``` + +The `table` URI uses the standard `s3://bucket/prefix` form; the endpoint and credentials come from `storage_options`. + +## 3. Read the table back + +```python title="delta_read.py" +from deltalake import DeltaTable + +back = DeltaTable("s3:///events", storage_options=storage_options).to_pandas() +print("read back:", back.shape[0], "rows") +print(back.sort_values("id").to_string(index=False)) +print("version:", DeltaTable("s3:///events", storage_options=storage_options).version()) +``` + +```text +read back: 3 rows + id name + 1 alpha + 2 beta + 3 gamma +version: 0 +``` + +Because the version is tracked in the transaction log, the same table supports time travel with `DeltaTable(..., version=N)` and appends that bump the version. + +## 4. Verify objects in RustFS + +List the table prefix: + +```bash +rc ls rustfs// -r +``` + +The first commit created the transaction log and one Parquet file: + +```text +events/_delta_log/00000000000000000000.json +events/part-00000-3859855e-45e4-4ae5-94ff-2d8eab5e7ebb-c000.snappy.parquet +``` + +Every new write adds a `NNNNNNNNNNNNNNNNNNNN.json` log entry and Parquet parts; readers replay the log to get a consistent snapshot. + +![Delta table files stored in the RustFS Console](./images/rustfs-delta-table.png) + +## 5. Stop or reset + +delta-rs holds no state of its own. To delete the table: + +```bash +rc rm rustfs//events/ --recursive --force +``` + +## Troubleshooting + +### `Import pyarrow failed` when writing a pandas DataFrame + +`write_deltalake` converts frames through Arrow. Install `pyarrow` alongside `deltalake` and `pandas`. + +### `Generic DeltaTable error: commit conflict` or rename errors on commit + +delta-rs commits by copying and renaming temporary objects, which requires atomic rename on the backend. For S3-compatible stores without copy-if-not-exists, set `AWS_S3_ALLOW_UNSAFE_RENAME: "true"` in `storage_options` — acceptable for a single writer, not for concurrent writers. + +### `Unknown lengthy error: AWS connectivity or endpoint errors` + +Confirm `AWS_ENDPOINT_URL` includes the scheme and that `AWS_ALLOW_HTTP` is `"true"` for plain-HTTP endpoints; without it the S3 client only speaks HTTPS. + +## Next steps + +- Review [S3 compatibility notes](/administration/protocols/s3) before adopting additional Delta clients. +- Create dedicated production credentials with [Access Key Management](/security-compliance/iam/access-token). +- Follow the [delta-rs usage documentation](https://delta-io.github.io/delta-rs/usage/writing/writing-to-s3/) for concurrent-writer setups with DynamoDB-backed commit coordination. diff --git a/content/fr/developer/integration/big-data/images/rustfs-airflow-object.png b/content/fr/developer/integration/big-data/images/rustfs-airflow-object.png new file mode 100644 index 00000000..9177d249 Binary files /dev/null and b/content/fr/developer/integration/big-data/images/rustfs-airflow-object.png differ diff --git a/content/fr/developer/integration/big-data/images/rustfs-delta-table.png b/content/fr/developer/integration/big-data/images/rustfs-delta-table.png new file mode 100644 index 00000000..093a5cf1 Binary files /dev/null and b/content/fr/developer/integration/big-data/images/rustfs-delta-table.png differ diff --git a/content/fr/developer/integration/big-data/images/rustfs-kafka-sink.png b/content/fr/developer/integration/big-data/images/rustfs-kafka-sink.png new file mode 100644 index 00000000..dd4b11b5 Binary files /dev/null and b/content/fr/developer/integration/big-data/images/rustfs-kafka-sink.png differ diff --git a/content/fr/developer/integration/big-data/index.md b/content/fr/developer/integration/big-data/index.md index 701a3710..2e820bff 100644 --- a/content/fr/developer/integration/big-data/index.md +++ b/content/fr/developer/integration/big-data/index.md @@ -8,6 +8,7 @@ Use **RustFS** as the object storage layer for data analytics systems that suppo ## Systems - [ClickHouse](./clickhouse.md) +- [Airflow](./airflow.md) - [Hudi](./hudi.md) - [Iceberg](./iceberg.md) - [PyIceberg](./pyiceberg.md) @@ -16,8 +17,10 @@ Use **RustFS** as the object storage layer for data analytics systems that suppo - [OpenDAL](./opendal.md) - [DuckDB](./duckdb.md) - [Doris](./doris.md) +- [Delta Lake](./delta-lake.md) - [lakeFS](./lakefs.md) - [InfluxDB](./influxdb.md) +- [Kafka](./kafka.md) - [Spark](./spark.md) - [Flink](./flink.md) - [Trino](./trino.md) diff --git a/content/fr/developer/integration/big-data/kafka.md b/content/fr/developer/integration/big-data/kafka.md new file mode 100644 index 00000000..ff1e92f9 --- /dev/null +++ b/content/fr/developer/integration/big-data/kafka.md @@ -0,0 +1,178 @@ +--- +title: "Kafka" +description: "Offload Kafka topic data to RustFS with the Kafka Connect S3 sink connector." +--- + +This guide connects [Apache Kafka](https://github.com/apache/kafka) — the distributed event streaming platform — to **RustFS** through the Kafka Connect S3 sink connector. You will run a KRaft broker and a Connect worker, deploy the S3 sink for a topic, and produce records that land as objects in a RustFS bucket. The workflow was verified with `apache/kafka:4.0.0` and `confluentinc/kafka-connect-s3` v10.5.25 against `rustfs/rustfs-x86-musl:v2.3.1`. + +Kafka's KIP-405 tiered storage needs a `RemoteLogStorageManager` plugin, and the S3 implementations in the ecosystem are vendor-proprietary. The Connect S3 sink is the open, self-hosted way to move topic data to S3-compatible storage and is the approach documented here. + +You need Docker. This deployment is intended for local integration testing, not production. + +## Architecture + +```mermaid +flowchart LR + Producer["Console producer"] -->|"records"| Broker["Kafka broker :9092"] + Broker -->|"consumer group"| Connect["Connect S3 sink"] + Connect -->|"batched objects"| RustFS["RustFS :9000"] +``` + +The sink task consumes a topic in a dedicated consumer group and writes record batches to the bucket, one object per `flush.size` records per partition. + +## 1. Run the broker + +Start a KRaft broker whose advertised listener is reachable from other containers: + +```bash +docker run -d --name kafka --hostname kafka --network oo-rustfs_default \ + -e CLUSTER_ID=5L6g3nShT-eMCtK--X86sw \ + -e KAFKA_NODE_ID=1 \ + -e KAFKA_PROCESS_ROLES=broker,controller \ + -e KAFKA_LISTENERS=PLAINTEXT://:9092,CONTROLLER://:9093 \ + -e KAFKA_ADVERTISED_LISTENERS=PLAINTEXT://kafka:9092 \ + -e KAFKA_CONTROLLER_LISTENER_NAMES=CONTROLLER \ + -e KAFKA_LISTENER_SECURITY_PROTOCOL_MAP=CONTROLLER:PLAINTEXT,PLAINTEXT:PLAINTEXT \ + -e KAFKA_CONTROLLER_QUORUM_VOTERS=1@kafka:9093 \ + -e KAFKA_OFFSETS_TOPIC_REPLICATION_FACTOR=1 \ + -e KAFKA_TRANSACTION_STATE_LOG_REPLICATION_FACTOR=1 \ + -e KAFKA_TRANSACTION_STATE_LOG_MIN_ISR=1 \ + apache/kafka:4.0.0 + +docker exec kafka /opt/kafka/bin/kafka-topics.sh \ + --bootstrap-server localhost:9092 \ + --create --topic rustfs-topic --partitions 1 --replication-factor 1 + +rc mb rustfs/kafka-demo +``` + +The image defaults to advertising `localhost:9092`, which only works from inside the broker container. The `KAFKA_ADVERTISED_LISTENERS` override is what makes Connect (and any remote client) able to reach the broker. + +## 2. Install the connector + +Download the Confluent Hub archive, which bundles the connector and its dependencies, and unpack it where the worker can see it: + +```bash +curl -Lo kafka-connect-s3.zip "https://hub-downloads.confluent.io/api/plugins/confluentinc/kafka-connect-s3/versions/10.5.25/confluentinc-kafka-connect-s3-10.5.25.zip" +unzip kafka-connect-s3.zip -d /opt/kafka-conn/plugins +``` + +## 3. Configure the worker and the sink + +Create the worker properties. The value converter must be `ByteArrayConverter` so records are written verbatim: + +```ini title="worker.properties" +bootstrap.servers=kafka:9092 +key.converter=org.apache.kafka.connect.storage.StringConverter +value.converter=org.apache.kafka.connect.converters.ByteArrayConverter +offset.storage.file.filename=/tmp/connect.offsets +offset.flush.interval.ms=5000 +plugin.path=/opt/kafka-conn-plugins +``` + +Create the sink connector configuration, replacing all connection placeholders: + +```ini title="rustfs-sink.properties" +name=rustfs-sink +connector.class=io.confluent.connect.s3.S3SinkConnector +tasks.max=1 +topics=rustfs-topic +s3.bucket.name=kafka-demo +s3.region=us-east-1 +store.url=http://:9000 +s3.path.style.access.enabled=true +flush.size=3 +storage.class=io.confluent.connect.s3.storage.S3Storage +format.class=io.confluent.connect.s3.format.bytearray.ByteArrayFormat +consumer.override.auto.offset.reset=earliest +``` + +Kafka 4.0 moved the class to `org.apache.kafka.connect.converters.ByteArrayConverter` — the old `storage` package path no longer resolves. + +## 4. Run the worker + +Run `connect-standalone` in the foreground so its logs go to `docker logs`, with the bucket credentials in the environment: + +```bash +docker run -d --name kafka-connect --hostname kafka-connect \ + --network oo-rustfs_default \ + -v /opt/kafka-conn/plugins:/opt/kafka-conn-plugins:ro \ + -v "$PWD/worker.properties":/etc/kafka/worker.properties:ro \ + -v "$PWD/rustfs-sink.properties":/etc/kafka/sink.properties:ro \ + -e AWS_ACCESS_KEY_ID= \ + -e AWS_SECRET_ACCESS_KEY= \ + apache/kafka:4.0.0 \ + /opt/kafka/bin/connect-standalone.sh /etc/kafka/worker.properties /etc/kafka/sink.properties +``` + +The worker is ready when the sink task claims the partition: + +```text +INFO [rustfs-sink|task-0] Assigned topic partitions: [rustfs-topic-0] +``` + +## 5. Produce records + +Send at least `flush.size` records so the connector completes a batch: + +```bash +docker exec kafka sh -c "printf 'msg-one\nmsg-two\nmsg-three\n' | \ + /opt/kafka/bin/kafka-console-producer.sh --bootstrap-server localhost:9092 --topic rustfs-topic" +``` + +After a few seconds the batch becomes an object in the bucket: + +```bash +rc ls rustfs/kafka-demo/ -r +rc cat rustfs/kafka-demo/topics/rustfs-topic/partition=0/rustfs-topic+0+0000000000.bin +``` + +```text +topics/rustfs-topic/partition=0/rustfs-topic+0+0000000000.bin +topics/rustfs-topic/partition=0/rustfs-topic+0+0000000003.bin +msg-one +msg-two +msg-three +``` + +The object name encodes topic, partition, and starting offset. Each subsequent batch of three records lands in the next object (`+0000000003.bin` and so on). + +![Kafka sink objects stored in the RustFS Console](./images/rustfs-kafka-sink.png) + +## 6. Stop or reset + +To tear down the demo while keeping the bucket objects: + +```bash +docker rm -f kafka-connect kafka +``` + +To delete the stored data: + +```bash +rc rm rustfs/kafka-demo/ --recursive --force +``` + +## Troubleshooting + +### `AdminClient ... Rebootstrapping with Cluster (id: null)` loops forever + +The broker advertises `localhost:9092`, so a remote client receives metadata pointing at itself. Set `KAFKA_ADVERTISED_LISTENERS=PLAINTEXT://kafka:9092` (and matching listener variables) as in step 1. + +### `Invalid schema type for ByteArrayConverter: STRING` + +The `ByteArrayFormat` writer only accepts raw bytes. Either switch the worker's `value.converter` to the ByteArray converter or choose a format class that matches the converter you use. + +### `Class org.apache.kafka.connect.storage.ByteArrayConverter could not be found` + +Kafka 4.0 moved the class to `org.apache.kafka.connect.converters.ByteArrayConverter`. Use the new package path in `worker.properties`. + +### The connector downloads but the plugin is not found + +The plain connector JAR from Maven lacks its dependencies. Use the Confluent Hub archive from step 2, which bundles the complete `lib/` directory, and make sure `plugin.path` points at the directory that contains the connector folder. + +## Next steps + +- Review [S3 compatibility notes](/administration/protocols/s3) before adopting additional connectors. +- Create dedicated production credentials with [Access Key Management](/security-compliance/iam/access-token). +- Follow the [Kafka Connect S3 sink documentation](https://docs.confluent.io/kafka-connect-s3/current/index.html) for Parquet and Avro formats, partitioning by time, and IAM-based credential chains. diff --git a/content/fr/developer/integration/big-data/meta.json b/content/fr/developer/integration/big-data/meta.json index 29684232..2374fa35 100644 --- a/content/fr/developer/integration/big-data/meta.json +++ b/content/fr/developer/integration/big-data/meta.json @@ -2,12 +2,15 @@ "title": "Data Analytics", "pages": [ "clickhouse", + "airflow", "duckdb", "doris", + "delta-lake", "flink", "hudi", "iceberg", "influxdb", + "kafka", "lakefs", "milvus", "mlflow", diff --git a/content/fr/developer/integration/devops/images/rustfs-opensearch-snapshot.png b/content/fr/developer/integration/devops/images/rustfs-opensearch-snapshot.png new file mode 100644 index 00000000..60b3bd3a Binary files /dev/null and b/content/fr/developer/integration/devops/images/rustfs-opensearch-snapshot.png differ diff --git a/content/fr/developer/integration/devops/index.md b/content/fr/developer/integration/devops/index.md index b66e4aa0..7ca9dbaf 100644 --- a/content/fr/developer/integration/devops/index.md +++ b/content/fr/developer/integration/devops/index.md @@ -8,6 +8,7 @@ Utilisez **RustFS** comme couche de stockage objet pour les plateformes DevOps e ## Plateformes et outils - [Elasticsearch](./elasticsearch.md) +- [OpenSearch](./opensearch.md) - [Gitea](./gitea.md) - [Jenkins](./jenkins.md) - [Terraform](./terraform.md) diff --git a/content/fr/developer/integration/devops/meta.json b/content/fr/developer/integration/devops/meta.json index 67c84c6e..55eada6a 100644 --- a/content/fr/developer/integration/devops/meta.json +++ b/content/fr/developer/integration/devops/meta.json @@ -2,6 +2,7 @@ "title": "DevOps", "pages": [ "elasticsearch", + "opensearch", "gitea", "jenkins", "terraform" diff --git a/content/fr/developer/integration/devops/opensearch.md b/content/fr/developer/integration/devops/opensearch.md new file mode 100644 index 00000000..9d44e530 --- /dev/null +++ b/content/fr/developer/integration/devops/opensearch.md @@ -0,0 +1,172 @@ +--- +title: "OpenSearch" +description: "Snapshot OpenSearch indices to RustFS with the repository-s3 plugin." +--- + +This guide connects [OpenSearch](https://github.com/opensearch-project/OpenSearch) — the open-source search and analytics suite derived from Elasticsearch — to **RustFS** through the `repository-s3` plugin. You will register an S3 snapshot repository backed by a RustFS bucket, take a snapshot of an index, and restore it. The workflow was verified with `opensearchproject/opensearch:3.8.0` and the bundled `repository-s3` plugin against `rustfs/rustfs-x86-musl:v2.3.1`. + +You need Docker, or an OpenSearch node where you can install plugins and edit configuration. This deployment is intended for local integration testing, not production. + +## Architecture + +```mermaid +flowchart LR + Client["REST client"] --> OS["OpenSearch :9200"] + OS -->|"snapshot files"| RustFS["RustFS :9000"] + RustFS -->|"restore"| OS +``` + +The `repository-s3` plugin writes snapshots as shard archives plus metadata blobs in the bucket. Registration is cluster-wide, so every node needs the plugin and the same client configuration. + +## 1. Run OpenSearch + +Start a single node with security disabled and a small heap: + +```bash +docker run -d --name opensearch --network oo-rustfs_default -p 9200:9200 \ + -e discovery.type=single-node \ + -e OPENSEARCH_JAVA_OPTS="-Xms512m -Xmx512m" \ + -e DISABLE_SECURITY_PLUGIN=true \ + opensearchproject/opensearch:3.8.0 +``` + +The node is ready when `curl http://localhost:9200` returns the cluster header (allow one to two minutes). + +## 2. Install the repository-s3 plugin + +The S3 repository plugin is not preloaded. Install it and restart the node: + +```bash +docker exec opensearch bin/opensearch-plugin install --batch repository-s3 +docker restart opensearch +``` + +## 3. Configure the S3 client + +Credentials are secure settings: they belong in the OpenSearch keystore, not in the repository request or `opensearch.yml`. Create the keystore entries, replacing all connection placeholders: + +```bash +docker exec opensearch sh -c \ + "printf '' | bin/opensearch-keystore create 2>/dev/null; \ + printf '' | bin/opensearch-keystore add -f -x s3.client.default.access_key; \ + printf '' | bin/opensearch-keystore add -f -x s3.client.default.secret_key" +``` + +Add the non-secure client settings to `config/opensearch.yml`: + +```yaml title="opensearch.yml" +network.host: 0.0.0.0 +plugins.security.disabled: true +s3.client.default.endpoint: http://:9000 +s3.client.default.protocol: http +s3.client.default.path_style_access: "true" +``` + +Restart the node once more so it reads both the keystore and the new settings: + +```bash +docker restart opensearch +``` + +Create the bucket while the node boots: + +```bash +rc mb rustfs/opensearch-snapshots +``` + +## 4. Register the repository and snapshot + +Create a test index with a document, then register the repository: + +```bash +curl -sX PUT http://localhost:9200/rustfs-demo -H "Content-Type: application/json" \ + -d '{"settings":{"number_of_shards":1}}' + +curl -sX PUT http://localhost:9200/rustfs-demo/_doc/1 -H "Content-Type: application/json" \ + -d '{"product":"rustfs","via":"opensearch-snapshot"}' + +curl -sX PUT "http://localhost:9200/_snapshot/rustfs-repo" -H "Content-Type: application/json" \ + -d '{"type":"s3","settings":{"bucket":"opensearch-snapshots","region":"us-east-1","server_side_encryption_type":"bucket_default"}}' +``` + +The `server_side_encryption_type: bucket_default` setting matters: without it the plugin requests SSE-S3, which a self-hosted RustFS without a server-side encryption master key rejects. + +Take a snapshot and wait for completion: + +```bash +curl -sX PUT "http://localhost:9200/_snapshot/rustfs-repo/snapshot-1?wait_for_completion=true" \ + -H "Content-Type: application/json" -d '{"indices":"rustfs-demo"}' +``` + +```text +{"snapshot":{"snapshot":"snapshot-1","state":"SUCCESS","indices":["rustfs-demo"],...}} +``` + +## 5. Verify objects and restore + +List the bucket: + +```bash +rc ls rustfs/opensearch-snapshots/ -r +``` + +```text +index-0 +index.latest +indices/5x1bwsWaSv2XINIbeoe-RQ/0/__GgxvoCw-TKuMMAYBq5Khag +indices/5x1bwsWaSv2XINIbeoe-RQ/0/snap-kRgFBiuyQIyPo3_-CMp7Hw.dat +meta-kRgFBiuyQIyPo3_-CMp7Hw.dat +snap-kRgFBiuyQIyPo3_-CMp7Hw.dat +``` + +Delete the index and restore it from the snapshot: + +```bash +curl -sX DELETE http://localhost:9200/rustfs-demo +curl -sX POST "http://localhost:9200/_snapshot/rustfs-repo/snapshot-1/_restore?wait_for_completion=true" +curl -s http://localhost:9200/rustfs-demo/_doc/1 +``` + +```text +{"_index":"rustfs-demo","_id":"1","found":true,"_source":{"product":"rustfs","via":"opensearch-snapshot"}} +``` + +![OpenSearch snapshot stored in the RustFS Console](./images/rustfs-opensearch-snapshot.png) + +## 6. Stop or reset + +To tear down the demo while keeping the bucket objects: + +```bash +docker rm -f opensearch +``` + +To delete the stored snapshots: + +```bash +rc rm rustfs/opensearch-snapshots/ --recursive --force +``` + +## Troubleshooting + +### `Setting [access_key] is insecure, but property [allow_insecure_settings] is not set` + +Inline credentials in the repository request are rejected. Store them in the keystore as shown in step 3 — `access_key` and `secret_key` are secure settings in OpenSearch. + +### `SSE-S3 requires RUSTFS_SSE_S3_MASTER_KEY ... (Status Code: 400)` + +The plugin encrypts uploads with SSE-S3 by default. Register the repository with `"server_side_encryption_type": "bucket_default"` so no encryption header is sent, as in step 4. + +### `unknown setting [s3.client.default.access_key]` at startup + +The settings reference the repository-s3 plugin. If the node fails to start with them present, the plugin is not installed in that container — repeat step 2 (a fresh container loses plugins installed with `docker exec`). + +### Repository verification fails with `path is not accessible` + +The node cannot reach the bucket: check that `s3.client.default.endpoint` is reachable from the container, `path_style_access` is `"true"`, and the keystore credentials were loaded (they are read at startup — restart after adding them). + +## Next steps + +- Review [S3 compatibility notes](/administration/protocols/s3) before adopting additional OpenSearch repositories. +- Create dedicated production credentials with [Access Key Management](/security-compliance/iam/access-token). +- Follow the [OpenSearch snapshots documentation](https://docs.opensearch.org/docs/latest/tuning-your-cluster/availability-and-recovery/snapshots/index/) to automate snapshots with Snapshot Management (SM) policies. diff --git a/content/fr/developer/integration/index.md b/content/fr/developer/integration/index.md index af4a9c16..d50853c1 100644 --- a/content/fr/developer/integration/index.md +++ b/content/fr/developer/integration/index.md @@ -7,14 +7,14 @@ Utilisez cette section pour connecter **RustFS** à des plateformes d'infrastruc ## Integration categories -- [Reverse Proxy](./reverse-proxy/index.md) couvre Nginx, Traefik, Caddy et HAProxy. +- [Reverse Proxy](./reverse-proxy/index.md) couvre Nginx, Traefik, Caddy, HAProxy et Envoy. - [Backup](./backup/index.md) couvre Kopia, Longhorn, Restic et Velero. -- [IA](./ai/index.md) couvre les plateformes d'IA incluant Ray. -- [Analyse de données](./big-data/index.md) couvre les systèmes d'analyse incluant ClickHouse, Doris, Hudi, Iceberg, lakeFS, Milvus, OpenDAL, Vitess et Zeppelin. +- [IA](./ai/index.md) couvre les plateformes d'IA incluant Ray et vLLM. +- [Analyse de données](./big-data/index.md) couvre les systèmes d'analyse incluant Airflow, ClickHouse, Delta Lake, Doris, Hudi, Iceberg, Kafka, lakeFS, Milvus, OpenDAL, Vitess et Zeppelin. - [Cloud Native](./cloud-native/index.md) couvre Cortex et Flux. -- [Observabilité](./observability/index.md) couvre les systèmes de télémétrie incluant Fluentd, OpenObserve, OpenTelemetry, Thanos et Tempo. -- [Autres](./others/index.md) couvre le SDK communautaire capo pour Python. +- [Observabilité](./observability/index.md) couvre les systèmes de télémétrie incluant Fluentd, GreptimeDB, Loki, OpenObserve, OpenTelemetry, Tempo, Thanos et VictoriaMetrics. +- [Autres](./others/index.md) couvre le SDK capo, rclone, JuiceFS, Nextcloud et tusd. - [Registre](./registry/index.md) couvre Harbor. -- [DevOps](./devops/index.md) couvre Elasticsearch, Gitea, Jenkins et Terraform. +- [DevOps](./devops/index.md) couvre Elasticsearch, Gitea, Jenkins, OpenSearch et Terraform. Chaque guide indique le point de terminaison RustFS et les exigences d'adressage à utiliser lors de la configuration du système intégré. \ No newline at end of file diff --git a/content/fr/developer/integration/observability/greptimedb.md b/content/fr/developer/integration/observability/greptimedb.md new file mode 100644 index 00000000..5b7ee690 --- /dev/null +++ b/content/fr/developer/integration/observability/greptimedb.md @@ -0,0 +1,129 @@ +--- +title: "GreptimeDB" +description: "Run GreptimeDB with RustFS as the S3-compatible object storage backend." +--- + +This guide connects [GreptimeDB](https://github.com/GreptimeTeam/greptimedb) — the open-source, cloud-native time-series database — to **RustFS** as its object storage backend. You will start a standalone instance with its `[storage]` section pointed at a RustFS bucket, write time-series rows through the SQL API, and confirm the Parquet files and manifests in the bucket. The workflow was verified with `greptime/greptimedb` (main, commit `179ff8e5`) against `rustfs/rustfs-x86-musl:v2.3.1`. + +You need Docker, or a local GreptimeDB binary. This deployment is intended for local integration testing, not production. + +## Architecture + +```mermaid +flowchart LR + SQL["SQL / Prometheus API"] --> DB["GreptimeDB"] + DB -->|"SST + manifests"| RustFS["RustFS :9000"] +``` + +GreptimeDB keeps its write-ahead log and recent data locally, then persists SSTables (Parquet) and table manifests to object storage. Pointing the storage backend at RustFS makes the bucket the durable home of all table data. + +## 1. Configure the storage backend + +Create the bucket and a config file with an S3 storage section, replacing all connection placeholders: + +```toml title="greptimedb.toml" +[storage] +type = "S3" +bucket = "" +root = "greptimedb" +access_key_id = "" +secret_access_key = "" +endpoint = "http://:9000" +region = "us-east-1" +``` + +GreptimeDB uses path-style requests for custom endpoints by default; virtual-hosted style must be opted into explicitly with `enable_virtual_host_style`, so no extra flag is needed for RustFS. + +## 2. Run GreptimeDB + +Start a standalone instance with the config file: + +```bash +docker run -d --name greptimedb --network oo-rustfs_default -p 4000:4000 -p 4002:4002 \ + -v "$PWD/greptimedb.toml":/etc/greptimedb/greptimedb.toml:ro \ + greptime/greptimedb:latest standalone start \ + --http-addr 0.0.0.0:4000 \ + --mysql-addr 0.0.0.0:4002 \ + --config-file /etc/greptimedb/greptimedb.toml +``` + +Port `4000` serves the HTTP SQL endpoint and `4002` the MySQL protocol. + +## 3. Write and query time series + +Create a table, insert rows, and read them back. The HTTP SQL endpoint takes form-encoded requests: + +```bash +curl -s -X POST "http://localhost:4000/v1/sql" \ + --data-urlencode "sql=CREATE TABLE rustfs_demo (host STRING, cpu DOUBLE, mem DOUBLE, ts TIMESTAMP TIME INDEX)" + +curl -s -X POST "http://localhost:4000/v1/sql" \ + --data-urlencode "sql=INSERT INTO rustfs_demo VALUES (\"node-1\", 0.31, 0.62, 1790681000000), (\"node-1\", 0.35, 0.63, 1790681060000), (\"node-2\", 0.51, 0.71, 1790681000000)" +``` + +```text +{"output":[{"affectedrows":3}],"execution_time_ms":2} +``` + +Query the rows back: + +```bash +curl -s -X POST "http://localhost:4000/v1/sql" \ + --data-urlencode "sql=SELECT * FROM rustfs_demo ORDER BY ts" +``` + +```text +{"output":[{"records":{"rows":[["node-2",0.51,0.71,1790681000000],["node-1",0.35,0.63,1790681060000]],"total_rows":2}}]} +``` + +## 4. Verify objects in RustFS + +List the bucket — after the memtable flushes, the bucket holds Parquet SSTables and JSON manifests: + +```bash +rc ls rustfs// -r +``` + +```text +greptimedb/data/greptime/public/1024/1024_0000000000/manifest/00000000000000000000.json +greptimedb/data/greptime/greptime_private/1025/1025_0000000000/b11e8b25-5763-4f05-bcab-b6ee0a756a69.parquet +greptimedb/data/greptime/greptime_private/1025/1025_0000000000/manifest/00000000000000000001.json +``` + +Each database gets a directory under `data/`, and per-region `manifest/*.json` files describe the SSTables GreptimeDB reads back during queries. + +![GreptimeDB data stored in the RustFS Console](./images/rustfs-greptimedb-data.png) + +## 5. Stop or reset + +To tear down the demo while keeping the bucket objects: + +```bash +docker rm -f greptimedb +``` + +To delete the stored data: + +```bash +rc rm rustfs// --recursive --force +``` + +## Troubleshooting + +### `Form requests must have Content-Type: application/x-www-form-urlencoded` + +The `/v1/sql` HTTP endpoint only accepts form-encoded bodies. Pass SQL with `curl --data-urlencode "sql=..."` (or `application/x-www-form-urlencoded`), not as a JSON body. + +### Bucket stays empty + +GreptimeDB flushes memtables to object storage asynchronously. Run a few more inserts and wait a few seconds, or trigger a manual flush, then list the bucket again. + +### Startup fails with an S3 error + +Confirm `endpoint` includes the scheme, the bucket exists, and `access_key_id`/`secret_access_key` match a RustFS access key. The `root` value is optional but keeps the table tree under a known prefix. + +## Next steps + +- Review [S3 compatibility notes](/administration/protocols/s3) before adopting additional GreptimeDB storage options. +- Create dedicated production credentials with [Access Key Management](/security-compliance/iam/access-token). +- Follow the [GreptimeDB configuration reference](https://docs.greptime.com/operational-guide/configure/configure-datanode/) to tune flush intervals and cache layers for production workloads. diff --git a/content/fr/developer/integration/observability/images/rustfs-greptimedb-data.png b/content/fr/developer/integration/observability/images/rustfs-greptimedb-data.png new file mode 100644 index 00000000..1874f210 Binary files /dev/null and b/content/fr/developer/integration/observability/images/rustfs-greptimedb-data.png differ diff --git a/content/fr/developer/integration/observability/images/rustfs-vm-backups.png b/content/fr/developer/integration/observability/images/rustfs-vm-backups.png new file mode 100644 index 00000000..a0bd0cee Binary files /dev/null and b/content/fr/developer/integration/observability/images/rustfs-vm-backups.png differ diff --git a/content/fr/developer/integration/observability/index.md b/content/fr/developer/integration/observability/index.md index 8850eba3..f798d27a 100644 --- a/content/fr/developer/integration/observability/index.md +++ b/content/fr/developer/integration/observability/index.md @@ -8,10 +8,12 @@ Utilisez **RustFS** comme couche de stockage objet pour les plateformes d'observ ## Plateformes - [Fluentd](./fluentd.md) +- [GreptimeDB](./greptimedb.md) - [OpenObserve](./openobserve.md) - [OpenTelemetry](./opentelemetry.md) - [Loki](./loki.md) - [Tempo](./tempo.md) - [Thanos](./thanos.md) +- [VictoriaMetrics](./victoriametrics.md) Conservez les données de télémétrie dans un bucket dédié et limitez les identifiants aux opérations de bucket requises. diff --git a/content/fr/developer/integration/observability/meta.json b/content/fr/developer/integration/observability/meta.json index 85efe3c7..a2527722 100644 --- a/content/fr/developer/integration/observability/meta.json +++ b/content/fr/developer/integration/observability/meta.json @@ -2,10 +2,12 @@ "title": "Observabilité", "pages": [ "fluentd", + "greptimedb", "loki", "openobserve", "opentelemetry", "tempo", - "thanos" + "thanos", + "victoriametrics" ] } diff --git a/content/fr/developer/integration/observability/victoriametrics.md b/content/fr/developer/integration/observability/victoriametrics.md new file mode 100644 index 00000000..efc799d9 --- /dev/null +++ b/content/fr/developer/integration/observability/victoriametrics.md @@ -0,0 +1,157 @@ +--- +title: "VictoriaMetrics" +description: "Back up VictoriaMetrics snapshots to RustFS with vmbackup." +--- + +This guide connects [VictoriaMetrics](https://github.com/VictoriaMetrics/VictoriaMetrics) — the Prometheus-compatible time-series database — to **RustFS** through `vmbackup` and `vmrestore`. You will run a single-node instance, import metrics, create an instant snapshot, back it up to a RustFS bucket, and restore the data into a fresh directory. The workflow was verified with `victoria-metrics`, `vmbackup`, and `vmrestore` v1.x images against `rustfs/rustfs-x86-musl:v2.3.1`. + +You need Docker. This deployment is intended for local integration testing, not production. + +## Architecture + +```mermaid +flowchart LR + Import["Prometheus import API"] --> VM["VictoriaMetrics :8428"] + VM -->|"instant snapshot"| Backup["vmbackup"] + Backup -->|"S3 API"| RustFS["RustFS :9000"] + RustFS -->|"restore"| Restore["vmrestore"] +``` + +`vmbackup` uploads a consistent point-in-time snapshot of the storage directory to any S3-compatible endpoint. `vmrestore` reverses the process, producing a data directory a VictoriaMetrics instance can open directly. + +## 1. Run VictoriaMetrics + +Create the bucket and start a single-node instance: + +```bash +rc mb rustfs/vm-backups + +docker run -d --name vm --network oo-rustfs_default -p 8428:8428 \ + -v vm-data:/storage \ + victoriametrics/victoria-metrics:latest \ + -storageDataPath=/storage -retentionPeriod=100y +``` + +## 2. Import metrics + +Write a couple of samples through the Prometheus import API: + +```bash +echo "vm_demo_metric 123" | curl -s --data-binary @- http://localhost:8428/api/v1/import/prometheus +echo "vm_demo_metric 456" | curl -s --data-binary @- http://localhost:8428/api/v1/import/prometheus +``` + +The endpoint answers `204 No Content`. Confirm the data is queryable: + +```bash +curl -s "http://localhost:8428/api/v1/export?match[]=vm_demo_metric" +``` + +```text +{"metric":{"__name__":"vm_demo_metric"},"values":[123,456],"timestamps":[1790680894604,1790680894619]} +``` + +## 3. Create a snapshot + +Ask VictoriaMetrics for a consistent snapshot: + +```bash +curl -s http://localhost:8428/snapshot/create +``` + +```text +{"status":"ok","snapshot":"20260929112134-18D9C6C76704F913"} +``` + +## 4. Back the snapshot up to RustFS + +Run `vmbackup` against the same storage volume, replacing the credential placeholders. The snapshot name comes from step 3: + +```bash +docker run --rm --network oo-rustfs_default \ + -e AWS_ACCESS_KEY_ID= \ + -e AWS_SECRET_ACCESS_KEY= \ + --volumes-from vm \ + victoriametrics/vmbackup:latest \ + -storageDataPath=/storage \ + -snapshotName=20260929112134-18D9C6C76704F913 \ + -dst=s3://vm-backups/demo \ + -customS3Endpoint=http://:9000 +``` + +```text +backup ... to S3{bucket: "vm-backups", dir: "demo/"} is complete; uploaded 760 bytes +``` + +`-customS3Endpoint` redirects the AWS SDK to RustFS; custom endpoints are addressed with path-style requests automatically. Set `AWS_EC2_METADATA_DISABLED=true` on hosts without an EC2 metadata service to skip credential lookup delays. + +## 5. Verify and restore + +List the bucket prefix: + +```bash +rc ls rustfs/vm-backups/demo/ +``` + +```text +backup_complete.ignore +backup_metadata.ignore +data/ +metadata/ +``` + +`backup_complete.ignore` marks a complete backup. Restore it into a fresh directory: + +```bash +docker run --rm --network oo-rustfs_default \ + -e AWS_ACCESS_KEY_ID= \ + -e AWS_SECRET_ACCESS_KEY= \ + -v /opt/vm-restore:/restore \ + victoriametrics/vmrestore:latest \ + -src=s3://vm-backups/demo \ + -storageDataPath=/restore \ + -customS3Endpoint=http://:9000 +``` + +```text +restored 760 bytes from backup in 0.055 seconds +``` + +The restored directory contains `data/`, `metadata/`, and a lock file — exactly what a VictoriaMetrics instance expects at `-storageDataPath`. + +![VictoriaMetrics backup stored in the RustFS Console](./images/rustfs-vm-backups.png) + +## 6. Stop or reset + +To tear down the demo while keeping the bucket objects: + +```bash +docker rm -f vm +docker volume rm vm-data +``` + +To delete the stored backups: + +```bash +rc rm rustfs/vm-backups/ --recursive --force +``` + +## Troubleshooting + +### `vmbackup` hangs at startup or fails to find credentials + +The AWS SDK probes the EC2 metadata service when environment credentials are absent. On machines without IMDS, export `AWS_EC2_METADATA_DISABLED=true` next to the key variables. + +### Backup parts re-upload on every run + +`vmbackup` performs incremental backups by comparing local and remote file hashes. Restoring to a fresh directory and running `vmbackup` from there re-uploads everything; keep the original data directory for incremental runs. + +### Query returns nothing right after import + +Imports are accepted asynchronously and the instant query endpoint can lag on a busy single node. Verify with `/api/v1/export` (or wait a few seconds) before creating the snapshot. + +## Next steps + +- Review [S3 compatibility notes](/administration/protocols/s3) before adopting additional VictoriaMetrics components. +- Create dedicated production credentials with [Access Key Management](/security-compliance/iam/access-token). +- Follow the [vmbackup documentation](https://docs.victoriametrics.com/vmbackup/) to schedule backups and prune old snapshots. diff --git a/content/fr/developer/integration/others/images/rustfs-juicefs-chunks.png b/content/fr/developer/integration/others/images/rustfs-juicefs-chunks.png new file mode 100644 index 00000000..737490b6 Binary files /dev/null and b/content/fr/developer/integration/others/images/rustfs-juicefs-chunks.png differ diff --git a/content/fr/developer/integration/others/images/rustfs-nextcloud-file.png b/content/fr/developer/integration/others/images/rustfs-nextcloud-file.png new file mode 100644 index 00000000..506e0d79 Binary files /dev/null and b/content/fr/developer/integration/others/images/rustfs-nextcloud-file.png differ diff --git a/content/fr/developer/integration/others/images/rustfs-rclone-sync.png b/content/fr/developer/integration/others/images/rustfs-rclone-sync.png new file mode 100644 index 00000000..dcf13712 Binary files /dev/null and b/content/fr/developer/integration/others/images/rustfs-rclone-sync.png differ diff --git a/content/fr/developer/integration/others/images/rustfs-tus-uploads.png b/content/fr/developer/integration/others/images/rustfs-tus-uploads.png new file mode 100644 index 00000000..e82a55e1 Binary files /dev/null and b/content/fr/developer/integration/others/images/rustfs-tus-uploads.png differ diff --git a/content/fr/developer/integration/others/index.md b/content/fr/developer/integration/others/index.md index 01f3b3a5..10cf9eeb 100644 --- a/content/fr/developer/integration/others/index.md +++ b/content/fr/developer/integration/others/index.md @@ -8,3 +8,7 @@ Guides d'intégration qui ne correspondent pas aux autres catégories. ## Guides - [capo (Python)](./capo.md) — connectez le SDK communautaire capo à RustFS avec des clients synchrones ou asynchrones. +rclone](./rclone.md) — sync, mount, and serve RustFS buckets from the command line. +- [tusd](./tusd.md) — receive resumable uploads into a RustFS bucket over the tus protocol. +- [JuiceFS](./juicefs.md) — mount a POSIX filesystem backed by a RustFS bucket. +- [Nextcloud](./nextcloud.md) — use RustFS as S3 external storage for Nextcloud files. diff --git a/content/fr/developer/integration/others/juicefs.md b/content/fr/developer/integration/others/juicefs.md new file mode 100644 index 00000000..428b08ce --- /dev/null +++ b/content/fr/developer/integration/others/juicefs.md @@ -0,0 +1,132 @@ +--- +title: "JuiceFS" +description: "Build a POSIX filesystem on RustFS with JuiceFS S3 object storage." +--- + +This guide connects [JuiceFS](https://github.com/juicedata/juicefs) — the cloud-native distributed POSIX filesystem — to **RustFS** as its object storage backend. You will format a volume whose data chunks live in a RustFS bucket, mount it locally, and read and write files through the mount. The workflow was verified with `juicefs v1.3.1` (community edition, SQLite metadata engine) against `rustfs/rustfs-x86-musl:v2.3.1`. + +You need the JuiceFS binary, a metadata engine, and FUSE (`fuse3` on Linux). This deployment is intended for local integration testing, not production. + +## Architecture + +```mermaid +flowchart LR + Mount["/mnt/jfs"] -->|"POSIX"| JuiceFS["JuiceFS client"] + JuiceFS -->|"metadata"| Meta["SQLite / Redis"] + JuiceFS -->|"data chunks"| RustFS["RustFS :9000"] +``` + +JuiceFS splits every file into chunks and stores them as objects under `chunks/` in the bucket, while the metadata engine tracks names, inodes, and layout. The filesystem behaves like a local disk but holds no data locally. + +## 1. Format the volume + +Create the bucket and format a JuiceFS volume backed by RustFS, replacing all connection placeholders. The bucket URL carries the endpoint, which selects path-style addressing: + +```bash +rc mb rustfs/jfs-demo + +juicefs format \ + --storage s3 \ + --bucket http://:9000/jfs-demo \ + --access-key \ + --secret-key \ + sqlite3:///opt/juicefs/jfs.db \ + rustfs-jfs +``` + +```text +Data use s3://:9000/jfs-demo/rustfs-jfs/ + OK, rustfs-jfs is ready +``` + +`sqlite3:///opt/juicefs/jfs.db` is the metadata engine for this test. In production, use Redis, MySQL, or PostgreSQL instead so multiple clients can mount the same volume. + +## 2. Mount the volume + +Mount the filesystem with the same metadata URL: + +```bash +mkdir -p /mnt/jfs +juicefs mount -d sqlite3:///opt/juicefs/jfs.db /mnt/jfs +``` + +```text +OK, rustfs-jfs is ready at /mnt/jfs +``` + +The `-d` flag runs the mount in the background. The volume is now a POSIX filesystem. + +## 3. Read and write files + +Use the mount like any other directory: + +```bash +echo "hello rustfs jfs" > /mnt/jfs/hello.txt +dd if=/dev/urandom of=/mnt/jfs/blob.bin bs=1M count=3 +mkdir -p /mnt/jfs/dir1 && echo nested > /mnt/jfs/dir1/nested.txt +cat /mnt/jfs/hello.txt +``` + +```text +hello rustfs jfs +``` + +Inspect the volume with `juicefs info`: + +```text +/mnt/jfs : + inode: 1 + files: 2 + dirs: 2 + length: 3.00 MiB +``` + +## 4. Verify chunks in RustFS + +List the bucket prefixes: + +```bash +rc ls rustfs/jfs-demo/ -r +``` + +Every file was split into content-addressed chunk objects: + +```text +rustfs-jfs/chunks/0/0/1_0_17 +rustfs-jfs/chunks/0/0/3_0_3145728 +rustfs-jfs/chunks/0/0/4_0_7 +``` + +The chunk name encodes the inode, chunk index, and size — for example `3_0_3145728` is the 3 MiB file written in step 3. + +![JuiceFS data chunks stored in the RustFS Console](./images/rustfs-juicefs-chunks.png) + +## 5. Stop or reset + +Unmount the volume, then optionally wipe the volume metadata and bucket data: + +```bash +juicefs umount /mnt/jfs +juicefs destroy --force sqlite3:///opt/juicefs/jfs.db rustfs-jfs +rc rm rustfs/jfs-demo/ --recursive --force +``` + +## Troubleshooting + +### `unknown option: --daemon` + +The background flag is a single dash: `juicefs mount -d`. Running without it keeps the mount in the foreground (useful for debugging). + +### `fusermount3: mount failed: Permission denied` + +Mounting requires the FUSE device. Inside a container, add `--device /dev/fuse --cap-add SYS_ADMIN` (or `--privileged`); on a host, install `fuse3` and confirm `/dev/fuse` exists. + +### Mount hangs or fails to reach storage + +The client must reach both the metadata engine and the bucket endpoint. Because the bucket URL embeds the endpoint, verify it from the mounting host with `curl` before formatting. + +## Next steps + +- Review [S3 compatibility notes](/administration/protocols/s3) before adopting additional JuiceFS storage backends. +- Create dedicated production credentials with [Access Key Management](/security-compliance/iam/access-token). +- Follow the [JuiceFS documentation](https://juicefs.com/docs/community/quick_start_guide/) to switch the metadata engine to Redis and mount the volume from multiple clients. diff --git a/content/fr/developer/integration/others/meta.json b/content/fr/developer/integration/others/meta.json index a879241e..50a3bf78 100644 --- a/content/fr/developer/integration/others/meta.json +++ b/content/fr/developer/integration/others/meta.json @@ -1,6 +1,10 @@ { "title": "Autres", "pages": [ - "capo" + "capo", + "juicefs", + "nextcloud", + "rclone", + "tusd" ] } diff --git a/content/fr/developer/integration/others/nextcloud.md b/content/fr/developer/integration/others/nextcloud.md new file mode 100644 index 00000000..f506d75c --- /dev/null +++ b/content/fr/developer/integration/others/nextcloud.md @@ -0,0 +1,145 @@ +--- +title: "Nextcloud" +description: "Use RustFS as S3 external storage for Nextcloud files." +--- + +This guide connects [Nextcloud](https://github.com/nextcloud/server) — the self-hosted content collaboration platform — to **RustFS** through its External Storage app with the S3 backend. You will enable `files_external`, mount a RustFS bucket into every user's files view, and upload a file through WebDAV that lands directly in the bucket. The workflow was verified with `nextcloud:32.0.15` (SQLite, single container) against `rustfs/rustfs-x86-musl:v2.3.1`. + +You need Docker, or an existing Nextcloud instance with `occ` access. This deployment is intended for local integration testing, not production. + +## Architecture + +```mermaid +flowchart LR + User["Browser / WebDAV"] --> Nextcloud["Nextcloud"] + Nextcloud -->|"files_external (S3)"| RustFS["RustFS :9000"] +``` + +Nextcloud proxies file operations on the mount point to the S3 backend. Objects are stored under their mount-relative paths, so the bucket mirrors the names users see. + +## 1. Install Nextcloud + +Run Nextcloud with an admin account, replacing all connection placeholders. SQLite keeps the test self-contained; use MariaDB or PostgreSQL in production: + +```bash +docker run -d --name nextcloud --network oo-rustfs_default -p 8080:80 \ + -e NEXTCLOUD_ADMIN_USER= \ + -e NEXTCLOUD_ADMIN_PASSWORD= \ + nextcloud:32.0.15 +``` + +If the web UI still shows the installer after startup, finish it manually: + +```bash +docker exec -u www-data nextcloud php occ maintenance:install \ + --admin-user --admin-password +``` + +## 2. Enable the External Storage app + +The `files_external` app ships with Nextcloud but starts disabled, and its `occ` commands only exist once the app is enabled: + +```bash +docker exec -u www-data nextcloud php occ app:enable files_external +``` + +```text +files_external 1.24.1 enabled +``` + +## 3. Mount the RustFS bucket + +Create an external storage of backend type `amazons3` with the `amazons3::accesskey` authentication backend. Replace all connection placeholders: + +```bash +docker exec -u www-data nextcloud php occ files_external:create \ + /rustfs amazons3 amazons3::accesskey \ + --user \ + --config bucket= \ + --config hostname= \ + --config port=9000 \ + --config use_ssl=false \ + --config use_path_style=true \ + --config key= \ + --config secret= +``` + +```text +Storage created with id 1 +``` + +The mount point `/rustfs` appears in the files view of the given user. `use_path_style=true` is required for a non-AWS endpoint. Check the connection before using it: + +```bash +docker exec -u www-data nextcloud php occ files_external:verify 1 +``` + +```text + - status: ok + - code: 0 +``` + +## 4. Upload a file and verify + +Upload through the WebDAV endpoint, which writes through the external storage: + +```bash +echo "nextcloud writes to rustfs" > /tmp/nc-demo.txt + +curl -u : \ + -T /tmp/nc-demo.txt \ + http://localhost:8080/remote.php/dav/files//rustfs/nc-demo.txt \ + -o /dev/null -w "%{http_code}\n" +``` + +```text +201 +``` + +Read it back through the same path, then confirm the object in RustFS: + +```bash +rc ls rustfs// -r +``` + +```text +[2026-09-29 13:42:48] 27 B nc-demo.txt +``` + +The object key equals the path inside the mount, so files uploaded through Nextcloud can also be read directly with any S3 client. + +![Nextcloud file stored in the RustFS Console](./images/rustfs-nextcloud-file.png) + +## 5. Stop or reset + +To remove the mount without touching the bucket: + +```bash +docker exec -u www-data nextcloud php occ files_external:delete 1 +``` + +To delete the bucket contents: + +```bash +rc rm rustfs// --recursive --force +``` + +## Troubleshooting + +### `There are no commands defined in the "files_external" namespace` + +The app is not enabled yet. Run `occ app:enable files_external` first; the `occ files_external:*` commands only register afterwards. + +### `Not enough arguments (missing: "authentication_backend")` + +`files_external:create` takes the storage backend and the authentication backend as two separate arguments: `amazons3 amazons3::accesskey`. The backend identifiers are listed by `occ files_external:backends`. + +### Mount shows but is empty, or uploads fail + +Confirm `hostname` is reachable from the Nextcloud container (use the container network name, not `localhost`), `use_path_style` is `true`, and the bucket exists. `occ files_external:verify ` reports the exact connection error. + +## Next steps + +- Review [S3 compatibility notes](/administration/protocols/s3) before adopting additional external storage backends. +- Create dedicated production credentials with [Access Key Management](/security-compliance/iam/access-token). +- Follow the [Nextcloud external storage documentation](https://docs.nextcloud.com/server/latest/admin_manual/configuration_files/external_storage_configuration_gui.html) to share the mount with groups and enable versioning. diff --git a/content/fr/developer/integration/others/rclone.md b/content/fr/developer/integration/others/rclone.md new file mode 100644 index 00000000..34f9fe4a --- /dev/null +++ b/content/fr/developer/integration/others/rclone.md @@ -0,0 +1,144 @@ +--- +title: "rclone" +description: "Sync, mount, and serve RustFS buckets with rclone over its S3-compatible API." +--- + +This guide connects [rclone](https://github.com/rclone/rclone) — the command-line tool for syncing files to and from cloud storage — to **RustFS** through its S3 backend. You will configure an S3 remote for RustFS, copy and sync files, read objects back, publish a bucket over HTTP with `rclone serve`, and mount the bucket as a local filesystem with `rclone mount`. The workflow was verified with `rclone v1.75.1` against `rustfs/rustfs-x86-musl:v2.3.1`. + +You need Docker, or a local rclone binary. This deployment is intended for local integration testing, not production. + +## Architecture + +```mermaid +flowchart LR + Files["Local files"] -->|"copy / sync"| Remote["rclone S3 remote"] + Remote -->|"S3 API"| RustFS["RustFS :9000"] + RustFS -->|"mount / serve"| Client["FUSE mount / HTTP clients"] +``` + +One remote definition drives every rclone command: data transfer, mounting, and serving all use the same S3 connection. + +## 1. Configure the remote + +Create an rclone config file with an S3 remote for RustFS, replacing all connection placeholders. The `Other` provider disables AWS-specific behavior, and path-style addressing is used automatically for custom endpoints: + +```ini title="rclone.conf" +[rustfs] +type = s3 +provider = Other +access_key_id = +secret_access_key = +endpoint = http://:9000 +region = us-east-1 +``` + +## 2. Copy and read objects + +Create the bucket and upload a directory with `rclone copy`: + +```bash +rc mb rustfs/rclone-demo +rclone copy /data rustfs:rclone-demo/seed +``` + +List and read back: + +```bash +rclone ls rustfs:rclone-demo/seed +rclone cat rustfs:rclone-demo/seed/hello.txt +``` + +```text + 3145728 blob.bin + 18 hello.txt +hello from rclone +``` + +`rclone lsd rustfs:` lists every bucket on the endpoint. + +## 3. Sync a directory + +`rclone sync` makes the destination identical to the source, including deletions. Remove a local file and sync: + +```bash +rm /data/hello.txt +rclone sync /data rustfs:rclone-demo/seed +rclone lsf rustfs:rclone-demo/seed +``` + +```text +blob.bin +``` + +`hello.txt` disappears from the bucket. Add `--dry-run` first to preview the changes without touching the bucket. + +## 4. Serve a bucket over HTTP + +Publish the bucket contents as an HTTP file server: + +```bash +rclone serve http --addr 0.0.0.0:8080 rustfs:rclone-demo/seed +``` + +Any HTTP client can now download objects: + +```bash +curl -s http://localhost:8080/blob.bin -o /dev/null -w "%{http_code} %{size_download} bytes\n" +``` + +```text +200 3145728 bytes +``` + +`rclone serve` also supports WebDAV, SFTP, and S3 endpoints over the same remote. + +## 5. Mount the bucket as a filesystem + +With FUSE available, mount the bucket locally and use it like a directory: + +```bash +rclone mount rustfs:rclone-demo /mnt/rclone --daemon +ls /mnt/rclone/seed +echo test > /mnt/rclone/write-test.txt +cat /mnt/rclone/write-test.txt +``` + +Files written through the mount appear in RustFS as regular objects: + +```bash +rc ls rustfs/rclone-demo/ -r +``` + +```text +[2026-09-29 11:12:41] 5 B write-test.txt +``` + +Unmount with `fusermount -u /mnt/rclone` when finished. + +## 6. Stop or reset + +rclone holds no server-side state. To delete the demo data: + +```bash +rclone purge rustfs:rclone-demo +``` + +## Troubleshooting + +### `Access Denied` or empty listings + +Confirm `endpoint` includes the scheme and that the key pair matches a RustFS access key. The `region` value is required by the S3 signer even though RustFS ignores it; keep `us-east-1`. + +### Mount fails with `fusermount3: mount failed: Permission denied` + +Mounting needs the FUSE device and elevated privileges. Inside a container, run with `--device /dev/fuse --cap-add SYS_ADMIN` — and use `--privileged` if the mount helper still fails. On a host, verify that `fuse3` is installed and `/dev/fuse` exists. + +### Sync deleted nothing on the destination + +`rclone copy` never deletes. Only `rclone sync` (or `rclone delete`) removes destination objects, and `--dry-run` is the safe way to preview either. + +## Next steps + +- Review [S3 compatibility notes](/administration/protocols/s3) before adopting additional rclone backends. +- Create dedicated production credentials with [Access Key Management](/security-compliance/iam/access-token). +- Follow the [rclone S3 documentation](https://rclone.org/s3/) for flags such as `--transfers`, bandwidth limits, and crypt overlays. diff --git a/content/fr/developer/integration/others/tusd.md b/content/fr/developer/integration/others/tusd.md new file mode 100644 index 00000000..99f3c588 --- /dev/null +++ b/content/fr/developer/integration/others/tusd.md @@ -0,0 +1,170 @@ +--- +title: "tusd" +description: "Receive resumable uploads into RustFS with the tusd server's S3 backend." +--- + +This guide connects [tusd](https://github.com/tus/tusd) — the official reference implementation of the tus resumable-upload protocol — to **RustFS** as its S3 storage backend. You will run tusd against a RustFS bucket, create an upload with the tus protocol, send the file in two chunks with an interruption in between, resume from the reported offset, and verify the assembled object in the bucket. The workflow was verified with `tusproject/tusd:v2.10.1` against `rustfs/rustfs-x86-musl:v2.3.1`. + +You need Docker, or a local tusd binary. This deployment is intended for local integration testing, not production. + +## Architecture + +```mermaid +flowchart LR + Client["tus client"] -->|"POST / PATCH / HEAD"| tusd["tusd :8080"] + tusd -->|"multipart upload"| RustFS["RustFS :9000"] +``` + +tusd stores each in-progress upload as S3 multipart parts in the bucket. A client that loses its connection asks the server for the last committed offset with `HEAD` and continues from there — the data already received is never sent twice. + +## 1. Run tusd + +Create the bucket and start tusd with the S3 backend, replacing all connection placeholders. The AWS region must be set even though RustFS ignores it: + +```bash +rc mb rustfs/tus-uploads + +docker run -d --name tusd --network oo-rustfs_default -p 8080:8080 \ + -e AWS_ACCESS_KEY_ID= \ + -e AWS_SECRET_ACCESS_KEY= \ + -e AWS_REGION=us-east-1 \ + tusproject/tusd:latest \ + -s3-bucket tus-uploads \ + -s3-endpoint http://:9000 +``` + +Check that the server is healthy: + +```bash +curl -s -o /dev/null -w "%{http_code}\n" http://localhost:8080/health +``` + +```text +200 +``` + +## 2. Create the upload + +Create a 6 MiB upload and read the `Location` header: + +```bash +curl -s -D - -o /dev/null -X POST http://localhost:8080/files/ \ + -H "Upload-Length: 6291456" -H "Tus-Resumable: 1.0.0" \ + | grep -i "^Location:" +``` + +```text +Location: http://localhost:8080/files/b7338250daa9...+NGVmNDRhZjEt... +``` + +The upload URL contains the file ID and a message-authentication tag. Strip the scheme and host before re-sending it (the server echoes whatever `Host` it received, which may not be reachable from your next client). + +## 3. Upload in chunks with an interruption + +Send the first 2.5 MB, then stop — this is the point where a mobile client would lose its connection: + +```bash +head -c 2500000 demo.bin > part1.bin + +curl -s -o /dev/null -w "%{http_code}\n" -X PATCH "http://localhost:8080${LOC}" \ + -H "Upload-Offset: 0" -H "Tus-Resumable: 1.0.0" \ + -H "Content-Type: application/offset+octet-stream" \ + --data-binary @part1.bin +``` + +```text +204 +``` + +Ask the server how much it actually has — this is the resumable-upload core: + +```bash +curl -s -X HEAD "http://localhost:8080${LOC}" \ + -H "Tus-Resumable: 1.0.0" -D - -o /dev/null | grep -i upload-offset +``` + +```text +Upload-Offset: 2500000 +``` + +## 4. Resume and finish + +Continue from offset 2500000 with the remaining bytes: + +```bash +tail -c 3791456 demo.bin > part2.bin + +curl -s -o /dev/null -w "%{http_code}\n" -X PATCH "http://localhost:8080${LOC}" \ + -H "Upload-Offset: 2500000" -H "Tus-Resumable: 1.0.0" \ + -H "Content-Type: application/offset+octet-stream" \ + --data-binary @part2.bin +``` + +```text +204 +``` + +Download the finished upload through tusd and compare checksums with the source: + +```bash +curl -s -o download.bin "http://localhost:8080${LOC}" +sha1sum demo.bin download.bin +``` + +```text +d9016032ced6c7515b67a0c556e006c4b25a5858 demo.bin +d9016032ced6c7515b67a0c556e006c4b25a5858 download.bin +``` + +## 5. Verify objects in RustFS + +List the bucket: + +```bash +rc ls rustfs/tus-uploads/ -r +``` + +The bucket holds the assembled object plus one `.info` metadata file per upload — both live and finished: + +```text +b7338250daa9a1a79c1343502b57b28f 6 MiB +b7338250daa9a1a79c1343502b57b28f.info 378 B +``` + +The object key is the upload ID, and the object body is the uploaded file byte-for-byte — so any S3 client can read completed uploads directly from the bucket. + +![tus uploads stored in the RustFS Console](./images/rustfs-tus-uploads.png) + +## 6. Stop or reset + +To tear down the server while keeping the bucket objects: + +```bash +docker rm -f tusd +``` + +To delete the stored uploads: + +```bash +rc rm rustfs/tus-uploads/ --recursive --force +``` + +## Troubleshooting + +### `CreateMultipartUpload ... A region must be set when sending requests to S3` + +tusd builds its S3 client from the AWS environment, and the region is mandatory for endpoint resolution. Export `AWS_REGION=us-east-1` next to the credentials, as in step 1. + +### `PATCH` returns `404` or connects to the wrong host + +The `Location` URL echoes the `Host` header of the creation request. When your client and the server use different hostnames (container name versus published port), strip the scheme and host from the URL and send the path to the address the client can reach. + +### Upload disappears after server restart + +The S3 backend keeps `.info` files in the bucket, so uploads survive restarts. If you run tusd against an empty bucket that another process prunes, the metadata is lost — protect the `tus-uploads` prefix from cleanup jobs. + +## Next steps + +- Review [S3 compatibility notes](/administration/protocols/s3) before adopting additional tusd backends. +- Create dedicated production credentials with [Access Key Management](/security-compliance/iam/access-token). +- Follow the [tus protocol documentation](https://tus.io/protocols/resumable-upload) for creation-with-upload, termination, and checksum extensions that tusd supports on top of the core protocol. diff --git a/content/fr/developer/integration/reverse-proxy/envoy.md b/content/fr/developer/integration/reverse-proxy/envoy.md new file mode 100644 index 00000000..dc635a48 --- /dev/null +++ b/content/fr/developer/integration/reverse-proxy/envoy.md @@ -0,0 +1,192 @@ +--- +title: "Envoy" +description: "Deploy RustFS behind Envoy with TLS-terminated routes for the S3 API and Console." +--- + +Use **Envoy** to terminate TLS and route separate hostnames to the RustFS S3 API and Console. This deployment runs Envoy and a single-node RustFS instance on one Docker network. You need Docker Engine, two DNS records, and a TLS certificate that covers both hostnames (the guide uses a self-signed certificate for testing). + +This guide uses these example hostnames: + +- `s3.example.com` for the S3 API +- `console.example.com` for the Console + +Replace them with hostnames that resolve to the Docker host. + +:::warning[Serve S3 from the root path] + +Do not publish the S3 API under a path such as `/s3/`. AWS Signature Version 4 includes the request path and host, so rewriting either value can invalidate signed requests. Envoy forwards the incoming `Host` header unchanged, which keeps signatures valid. + +::: + +## 1. Create the deployment directories + +Create directories for the Envoy configuration and TLS certificate: + +```bash +mkdir -p rustfs-envoy/certs +cd rustfs-envoy +``` + +For local testing, generate a self-signed certificate covering both hostnames: + +```bash +openssl req -x509 -newkey rsa:2048 -nodes \ + -keyout certs/privkey.pem -out certs/fullchain.pem -days 30 \ + -subj "/CN=*.example.com" \ + -addext "subjectAltName=DNS:s3.example.com,DNS:console.example.com" +chmod 644 certs/privkey.pem +``` + +Make the certificate readable by the non-root user the Envoy image runs as — a `600` private key produces a misleading `Failed to load incomplete private key` error. + +## 2. Configure Envoy + +Create the configuration with an HTTPS listener and two virtual hosts. Route timeouts are disabled (`timeout: 0s`) so long streaming S3 uploads are not cut off: + +```yaml title="envoy.yaml" +static_resources: + listeners: + - name: https + address: {socket_address: {address: 0.0.0.0, port_value: 8443}} + filter_chains: + - transport_socket: + name: envoy.transport_sockets.tls + typed_config: + "@type": type.googleapis.com/envoy.extensions.transport_sockets.tls.v3.DownstreamTlsContext + common_tls_context: + tls_certificates: + - certificate_chain: {filename: /certs/fullchain.pem} + private_key: {filename: /certs/privkey.pem} + filters: + - name: envoy.filters.network.http_connection_manager + typed_config: + "@type": type.googleapis.com/envoy.extensions.filters.network.http_connection_manager.v3.HttpConnectionManager + stat_prefix: rustfs_https + route_config: + virtual_hosts: + - name: s3 + domains: ["s3.example.com", "s3.example.com:*"] + routes: + - match: {prefix: "/"} + route: {cluster: rustfs_s3, timeout: 0s} + - name: console + domains: ["console.example.com", "console.example.com:*"] + routes: + - match: {prefix: "/"} + route: {cluster: rustfs_console, timeout: 0s} + http_filters: + - name: envoy.filters.http.router + typed_config: + "@type": type.googleapis.com/envoy.extensions.filters.http.router.v3.Router + clusters: + - name: rustfs_s3 + connect_timeout: 5s + type: STRICT_DNS + lb_policy: ROUND_ROBIN + load_assignment: + cluster_name: rustfs_s3 + endpoints: + - lb_endpoints: + - endpoint: {address: {socket_address: {address: rustfs, port_value: 9000}}} + - name: rustfs_console + connect_timeout: 5s + type: STRICT_DNS + lb_policy: ROUND_ROBIN + load_assignment: + cluster_name: rustfs_console + endpoints: + - lb_endpoints: + - endpoint: {address: {socket_address: {address: rustfs, port_value: 9001}}} +``` + +The `host:*` domain entries matter: clients send `Host: s3.example.com:8443` on non-standard ports, and Envoy matches the authority including the port. + +## 3. Start Envoy + +Run Envoy on the same Docker network as RustFS, publishing only the proxy port: + +```bash +docker run -d --name envoy --network oo-rustfs_default -p 8443:8443 \ + -v "$PWD/envoy.yaml":/envoy.yaml:ro \ + -v "$PWD/certs":/certs:ro \ + envoyproxy/envoy:v1.34-latest -c /envoy.yaml +``` + +## 4. Verify both endpoints + +Point the example hostnames at the proxy with `curl --resolve` (in production, DNS does this): + +```bash +curl -sk --resolve s3.example.com:8443:127.0.0.1 \ + https://s3.example.com:8443/health/ready -o /dev/null -w "s3 api: %{http_code}\n" + +curl -sk --resolve console.example.com:8443:127.0.0.1 \ + https://console.example.com:8443/rustfs/console/ -o /dev/null -w "console: %{http_code}\n" +``` + +```text +s3 api: 200 +console: 200 +``` + +`-k` skips certificate validation because the certificate is self-signed; with a trusted certificate, drop it. + +## 5. Send signed S3 requests through Envoy + +Point any S3 client at the proxy as if it were RustFS. Configure the client with `https://s3.example.com:8443` as the endpoint and path-style addressing; when the proxy certificate is trusted, signed AWS Signature Version 4 requests pass through unchanged. For a quick test over plain HTTP, add an HTTP listener on port 8080 with the same virtual-host routing as the HTTPS listener, then use the endpoint `http://s3.example.com:8080`: + +```bash +rc alias set rustfs-envoy http://s3.example.com:8080 +rc ls rustfs-envoy/rclone-demo/ +``` + +```text +[ ] 0B seed/ +[2026-09-29 11:12:41] 5 B write-test.txt +``` + +The signatures validate because Envoy forwards the original `Host` header to RustFS. + +## Multi-node backends + +For a distributed RustFS deployment, add every node to the S3 cluster: + +```yaml title="envoy.yaml" + - name: rustfs_s3 + connect_timeout: 5s + type: STRICT_DNS + lb_policy: ROUND_ROBIN + load_assignment: + cluster_name: rustfs_s3 + endpoints: + - lb_endpoints: + - endpoint: {address: {socket_address: {address: node1, port_value: 9000}}} + - lb_endpoints: + - endpoint: {address: {socket_address: {address: node2, port_value: 9000}}} + - lb_endpoints: + - endpoint: {address: {socket_address: {address: node3, port_value: 9000}}} + - lb_endpoints: + - endpoint: {address: {socket_address: {address: node4, port_value: 9000}}} +``` + +The Console cluster follows the same pattern on port `9001`. + +## Troubleshooting + +### `Failed to load incomplete private key from path` + +The Envoy container runs as a non-root user and cannot read a `600` root-owned key. `chmod 644` the key files (or chown them to the container user, UID `1001` in the official image). + +### Routes return `404` with the correct hostnames + +Envoy matches the authority including the port. Add the `host:*` variants to each virtual host's `domains` list, as in the configuration above. + +### `Access Denied` from RustFS on proxied requests + +Confirm the proxy is not rewriting the path or the `Host` header. Signed requests must reach RustFS with the host the client signed for. + +## Next steps + +- [Configure an S3 client](/developer/examples/aws-cli) +- [Enable virtual-hosted-style bucket URLs](/integration/virtual) +- [Review health and readiness endpoints](/operations/status-check) diff --git a/content/fr/developer/integration/reverse-proxy/index.md b/content/fr/developer/integration/reverse-proxy/index.md index eff7324d..397cdad6 100644 --- a/content/fr/developer/integration/reverse-proxy/index.md +++ b/content/fr/developer/integration/reverse-proxy/index.md @@ -13,6 +13,7 @@ We recommend using separate hostnames for the S3 API on port `9000` and the Cons - [Traefik](./traefik.md) - [Caddy](./caddy.md) - [HAProxy](./haproxy.md) +- [Envoy](./envoy.md) - [Apache HTTP Server](./httpd.md) ## Related configuration diff --git a/content/fr/developer/integration/reverse-proxy/meta.json b/content/fr/developer/integration/reverse-proxy/meta.json index b5a2d6b5..4ca6d8b2 100644 --- a/content/fr/developer/integration/reverse-proxy/meta.json +++ b/content/fr/developer/integration/reverse-proxy/meta.json @@ -4,7 +4,8 @@ "nginx", "traefik", "caddy", + "envoy", "haproxy", "httpd" ] -} \ No newline at end of file +} diff --git a/content/ja/developer/integration/ai/images/rustfs-vllm-models.png b/content/ja/developer/integration/ai/images/rustfs-vllm-models.png new file mode 100644 index 00000000..baa18dc1 Binary files /dev/null and b/content/ja/developer/integration/ai/images/rustfs-vllm-models.png differ diff --git a/content/ja/developer/integration/ai/index.md b/content/ja/developer/integration/ai/index.md index fb36b0fc..fa215ed6 100644 --- a/content/ja/developer/integration/ai/index.md +++ b/content/ja/developer/integration/ai/index.md @@ -8,5 +8,6 @@ S3 互換エンドポイントをサポートする AI・機械学習プラッ ## プラットフォーム - [Ray](./ray.md) +- [vLLM](./vllm.md) トレーニングデータセットとチェックポイントは専用バケットに保存し、必要なバケット操作のみに権限が絞られた認証情報を使用してください。 diff --git a/content/ja/developer/integration/ai/meta.json b/content/ja/developer/integration/ai/meta.json index 573fc520..b30f9e0f 100644 --- a/content/ja/developer/integration/ai/meta.json +++ b/content/ja/developer/integration/ai/meta.json @@ -1,6 +1,7 @@ { "title": "AI", "pages": [ - "ray" + "ray", + "vllm" ] } diff --git a/content/ja/developer/integration/ai/vllm.md b/content/ja/developer/integration/ai/vllm.md new file mode 100644 index 00000000..c471096c --- /dev/null +++ b/content/ja/developer/integration/ai/vllm.md @@ -0,0 +1,154 @@ +--- +title: "vLLM" +description: "Serve LLM inference with vLLM loading model weights stored in RustFS." +--- + +This guide connects [vLLM](https://github.com/vllm-project/vllm) — the high-throughput LLM inference engine — to **RustFS** as its model-weight store. You will upload a model into a RustFS bucket, expose the bucket to the vLLM host through an rclone mount, and serve the model with the OpenAI-compatible API. The workflow was verified with `vllm/vllm-openai-cpu` (vLLM 0.30.0) serving `facebook/opt-125m` from a RustFS bucket backed by `rustfs/rustfs-x86-musl:v2.3.1`, on a CPU-only host. + +You need Docker and an rclone binary on the host that runs vLLM. This deployment is intended for local integration testing, not production. + +## Architecture + +```mermaid +flowchart LR + Upload["rclone copy"] -->|"weights"| RustFS["RustFS :9000"] + RustFS -->|"rclone mount"| Mount["/mnt/vllm-models"] + Mount -->|"weight load"| vLLM["vLLM :8000"] + Client["OpenAI SDK / curl"] -->|"completions"| vLLM +``` + +The bucket is the single copy of the model. Hosts that serve the model mount the bucket read-only, so every node pulls weights from RustFS and no local model store exists to drift. + +:::note[Why an rclone mount] + +vLLM 0.30 loads `s3://` model paths through the RunAI model streamer, whose ranged reads currently fail against custom S3 endpoints such as RustFS (the loader errors with `File access error` on any non-zero offset). Mounting the bucket as a filesystem is the verified way to keep the weights in RustFS while vLLM reads them as local files. + +::: + +## 1. Upload the model to RustFS + +Create the bucket and copy model weights into it, replacing all connection placeholders: + +```ini title="rclone.conf" +[rustfs] +type = s3 +provider = Other +access_key_id = +secret_access_key = +endpoint = http://:9000 +region = us-east-1 +``` + +```bash +rc mb rustfs/vllm-models +rclone copy ./opt-125m rustfs:vllm-models/opt-125m --transfers 4 +``` + +Any Hugging Face layout works — `config.json`, the tokenizer files, and the weight files (`model.safetensors` or `pytorch_model.bin`). Keep one model per prefix so several models can share the bucket. + +## 2. Mount the bucket on the vLLM host + +On the machine that runs vLLM, mount the bucket read-only for clients with `--allow-other`: + +```bash +mkdir -p /mnt/vllm-models +rclone mount rustfs:vllm-models /mnt/vllm-models \ + --allow-other --daemon +ls /mnt/vllm-models/opt-125m/ +``` + +```text +config.json merges.txt model.safetensors tokenizer.json vocab.json +``` + +## 3. Run vLLM + +Start the CPU image against the mounted weights: + +```bash +docker run -d --name vllm -p 8000:8000 --shm-size=2g \ + -v /mnt/vllm-models:/models:ro \ + vllm/vllm-openai-cpu:latest \ + --model /models/opt-125m --served-model-name opt-125m \ + --dtype float32 --max-model-len 256 --gpu-memory-utilization 0.15 +``` + +vLLM reads the weights through the mount — the container stays stateless and the model lives in RustFS. Wait for the server to come up: + +```bash +curl -s http://localhost:8000/v1/models | head -c 200 +``` + +```text +{"object":"list","data":[{"id":"opt-125m","object":"model","created":...,"root":"/models/opt-125m",...}]} +``` + +`--gpu-memory-utilization` controls the fraction of RAM reserved for the KV cache on the CPU backend; lower it on small hosts. `--dtype float32` matches what the CPU attention kernels support for this model. + +## 4. Run inference + +Send an OpenAI-compatible completion request: + +```bash +curl -s http://localhost:8000/v1/completions \ + -H "Content-Type: application/json" \ + -d '{"model": "opt-125m", "prompt": "RustFS is", "max_tokens": 12, "temperature": 0}' +``` + +```json +{"id":"cmpl-...","object":"text_completion","model":"opt-125m", + "choices":[{"index":0,"text":" a great tool for building your own server. It's a", + "finish_reason":"length",...}]} +``` + +The request is standard OpenAI schema, so the Python client works unchanged: + +```python +from openai import OpenAI + +client = OpenAI(base_url="http://localhost:8000/v1", api_key="EMPTY") +print(client.completions.create( + model="opt-125m", prompt="RustFS is", max_tokens=12, temperature=0, +).choices[0].text) +``` + +![vLLM model weights stored in the RustFS Console](./images/rustfs-vllm-models.png) + +## 5. Stop or reset + +To tear down the demo while keeping the bucket objects: + +```bash +docker rm -f vllm +fusermount -u /mnt/vllm-models +``` + +To delete the stored model: + +```bash +rclone purge rustfs:vllm-models +``` + +## Troubleshooting + +### `Cannot find any model weights with /models/...` + +The mount had a stale directory cache or the weight files never made it to the bucket. Re-run `rclone copy` and confirm the files through the mount with `ls` before starting vLLM. A short `--dir-cache-time` (for example `10s`) helps while you iterate. + +### `Unsupported CPU attention configuration: head_dim=...` + +vLLM's CPU kernels support a fixed set of head dimensions. Tiny test models such as `hf-internal-testing/tiny-random-*` use exotic shapes that fail at request time — use a real small model such as `facebook/opt-125m`. + +### `Insufficient space in /dev/shm` + +vLLM's CPU engine exchanges tensors through shared memory. Run the container with `--shm-size=2g` (or `--ipc=host`). + +### Server exits with `Available memory on node 0 ... is less than desired CPU memory utilization` + +The default KV-cache reservation is 90% of system RAM. Lower it with `--gpu-memory-utilization 0.15` (the flag applies to the CPU backend as a memory fraction despite its name). + +## Next steps + +- Review [S3 compatibility notes](/administration/protocols/s3) before adopting additional serving setups. +- Create dedicated production credentials with [Access Key Management](/security-compliance/iam/access-token). +- Follow the [vLLM documentation](https://docs.vllm.ai/en/latest/serving/openai_compatible_server.html) for chat templates, tensor parallelism, and quantized weights on top of the same bucket-backed model store. diff --git a/content/ja/developer/integration/big-data/airflow.md b/content/ja/developer/integration/big-data/airflow.md new file mode 100644 index 00000000..7f55e219 --- /dev/null +++ b/content/ja/developer/integration/big-data/airflow.md @@ -0,0 +1,170 @@ +--- +title: "Airflow" +description: "Move data between Airflow DAGs and RustFS with the Amazon S3 provider." +--- + +This guide connects [Apache Airflow](https://github.com/apache/airflow) — the workflow orchestration platform — to **RustFS** through the Amazon S3 provider's hooks, operators, and sensors. You will register a custom-endpoint connection, run a DAG that writes an object to a RustFS bucket, waits for a key with `S3KeySensor`, and reads the object back with `S3Hook`. The workflow was verified with `apache/airflow:3.3.2` (standalone, SequentialExecutor) and the `apache-airflow-providers-amazon` provider against `rustfs/rustfs-x86-musl:v2.3.1`. + +You need Docker, or an existing Airflow installation. This deployment is intended for local integration testing, not production. + +## Architecture + +```mermaid +flowchart LR + Scheduler["Airflow scheduler"] -->|"tasks"| Hook["S3Hook / operators"] + Hook -->|"S3 API"| RustFS["RustFS :9000"] + Sensor["S3KeySensor"] -->|"poll key"| RustFS +``` + +Every S3 interaction inside a DAG goes through the provider's S3 client, pointed at RustFS by the connection's `endpoint_url`. Operators, sensors, and hooks share the same connection object. + +## 1. Run Airflow + +Start a standalone instance with examples disabled, and create the demo bucket: + +```bash +docker run -d --name airflow --network oo-rustfs_default -p 8080:8080 \ + -e AIRFLOW__CORE__LOAD_EXAMPLES=False \ + -v "$PWD/dags":/opt/airflow/dags \ + apache/airflow:3.3.2 standalone + +rc mb rustfs/airflow-demo +``` + +The image ships with all providers preinstalled, including `apache-airflow-providers-amazon`. + +## 2. Register the RustFS connection + +The S3 provider reads its endpoint from the connection's extra field. Replace all connection placeholders: + +```bash +docker exec airflow airflow connections add rustfs \ + --conn-type aws \ + --conn-extra '{"endpoint_url": "http://:9000", "region_name": "us-east-1", "aws_access_key_id": "", "aws_secret_access_key": ""}' +``` + +The keys `aws_access_key_id` and `aws_secret_access_key` inside `--conn-extra` supply credentials; `endpoint_url` redirects the boto3 client from AWS to RustFS. + +## 3. Write the DAG + +The DAG writes an object with an operator, waits for the key with a sensor, and reads it back with the hook: + +```python title="rustfs_demo.py" +import datetime + +from airflow.providers.amazon.aws.hooks.s3 import S3Hook +from airflow.providers.amazon.aws.operators.s3 import S3CreateObjectOperator +from airflow.providers.amazon.aws.sensors.s3 import S3KeySensor +from airflow.sdk import dag, task + +@dag( + schedule=None, + start_date=datetime.datetime(2026, 1, 1), + catchup=False, + tags=["rustfs"], +) +def rustfs_demo(): + create = S3CreateObjectOperator( + task_id="write_object", + s3_bucket="airflow-demo", + s3_key="dags/airflow-put.txt", + data="written by airflow to rustfs", + aws_conn_id="rustfs", + replace=True, + ) + + wait = S3KeySensor( + task_id="wait_for_object", + bucket_key="dags/airflow-put.txt", + bucket_name="airflow-demo", + aws_conn_id="rustfs", + timeout=120, + poke_interval=10, + mode="reschedule", + ) + + @task + def read_object(): + hook = S3Hook(aws_conn_id="rustfs") + body = hook.read_key(key="dags/airflow-put.txt", bucket_name="airflow-demo") + print("read back:", body) + assert body == "written by airflow to rustfs" + + create >> [wait, read_object()] + +rustfs_demo() +``` + +Note the import paths: `S3CreateObjectOperator` lives in the `operators` module while `S3KeySensor` lives in the `sensors` module — importing both from one place fails. + +## 4. Unpause and trigger + +New DAGs start paused, and a trigger fired while paused stays queued forever. Unpause first, then trigger: + +```bash +docker exec airflow airflow dags unpause rustfs_demo +docker exec airflow airflow dags trigger rustfs_demo +``` + +Watch the run finish: + +```bash +docker exec airflow airflow dags list-runs rustfs_demo | head -3 +``` + +```text +dag_id run_id state +rustfs_demo manual__2026-09-29T13:15:57.332688+00:00 success +``` + +All three tasks succeed: `write_object`, `wait_for_object`, and `read_object`. + +## 5. Verify objects in RustFS + +List the bucket prefix: + +```bash +rc ls rustfs/airflow-demo/ -r +rc cat rustfs/airflow-demo/dags/airflow-put.txt +``` + +```text +[2026-09-29 13:16:01] 28 B dags/airflow-put.txt +written by airflow to rustfs +``` + +![Airflow object stored in the RustFS Console](./images/rustfs-airflow-object.png) + +## 6. Stop or reset + +To tear down the demo while keeping the bucket objects: + +```bash +docker rm -f airflow +``` + +To delete the stored data: + +```bash +rc rm rustfs/airflow-demo/ --recursive --force +``` + +## Troubleshooting + +### Dag runs stay `queued` after triggering + +The DAG is paused. New DAGs are paused by default in Airflow 3, and runs triggered in that state never execute. Run `airflow dags unpause rustfs_demo`; queued runs then start on their own. + +### `cannot import name 'S3KeySensor' from 'airflow.providers.amazon.aws.operators.s3'` + +The sensor lives in a separate module: `from airflow.providers.amazon.aws.sensors.s3 import S3KeySensor`. + +### Tasks fail with connection errors + +The `endpoint_url` must be reachable from the Airflow container — use the Docker network hostname for RustFS, not `localhost`. Airflow 3 serves its health endpoint under `/api/v2/monitor/health` if you need to check component status. + +## Next steps + +- Review [S3 compatibility notes](/administration/protocols/s3) before adopting additional provider hooks. +- Create dedicated production credentials with [Access Key Management](/security-compliance/iam/access-token). +- Follow the [Amazon provider documentation](https://airflow.apache.org/docs/apache-airflow-providers-amazon/stable/index.html) for transfer operators such as `S3ToLocalFilesystemOperator` and `LocalFilesystemToS3Operator`. diff --git a/content/ja/developer/integration/big-data/delta-lake.md b/content/ja/developer/integration/big-data/delta-lake.md new file mode 100644 index 00000000..b1bfae56 --- /dev/null +++ b/content/ja/developer/integration/big-data/delta-lake.md @@ -0,0 +1,125 @@ +--- +title: "Delta Lake" +description: "Write and read Delta tables on RustFS with delta-rs." +--- + +This guide connects [Delta Lake](https://github.com/delta-io/delta) — the open-source lakehouse table format — to **RustFS** through delta-rs, the Rust-native Delta implementation. You will write a Delta table to a RustFS bucket from Python, read it back with ACID transaction history, and confirm the `_delta_log` and Parquet files in the bucket. The workflow was verified with the `deltalake` Python package (delta-rs) and pandas against `rustfs/rustfs-x86-musl:v2.3.1`. + +You need Python 3.9 or newer. This deployment is intended for local integration testing, not production. + +## Architecture + +```mermaid +flowchart LR + DF["pandas DataFrame"] -->|"write_deltalake"| deltaRS["delta-rs"] + deltaRS -->|"Parquet + _delta_log"| RustFS["RustFS :9000"] + Query["DeltaTable"] -->|"read / time travel"| RustFS +``` + +delta-rs stores each table as Parquet files plus a transaction log (`_delta_log/`). All I/O goes through the `object_store` crate, configured with the same AWS environment variables as other S3 clients. + +## 1. Install the client + +```bash +pip install deltalake pandas pyarrow +``` + +`pyarrow` is required to convert pandas frames into Delta-compatible record batches. + +## 2. Write a Delta table + +Create the bucket and write a table, replacing all connection placeholders. `AWS_S3_ALLOW_UNSAFE_RENAME` is needed because RustFS does not provide copy-if-not-exists, which delta-rs otherwise uses for commit conflicts: + +```python title="delta_s3.py" +import pandas as pd +from deltalake import DeltaTable, write_deltalake + +storage_options = { + "AWS_ENDPOINT_URL": "http://:9000", + "AWS_ACCESS_KEY_ID": "", + "AWS_SECRET_ACCESS_KEY": "", + "AWS_REGION": "us-east-1", + "AWS_ALLOW_HTTP": "true", + "AWS_S3_ALLOW_UNSAFE_RENAME": "true", +} + +table = "s3:///events" +df = pd.DataFrame({"id": [1, 2, 3], "name": ["alpha", "beta", "gamma"]}) +write_deltalake(table, df, storage_options=storage_options) +print("written:", df.shape[0], "rows") +``` + +```text +written: 3 rows +``` + +The `table` URI uses the standard `s3://bucket/prefix` form; the endpoint and credentials come from `storage_options`. + +## 3. Read the table back + +```python title="delta_read.py" +from deltalake import DeltaTable + +back = DeltaTable("s3:///events", storage_options=storage_options).to_pandas() +print("read back:", back.shape[0], "rows") +print(back.sort_values("id").to_string(index=False)) +print("version:", DeltaTable("s3:///events", storage_options=storage_options).version()) +``` + +```text +read back: 3 rows + id name + 1 alpha + 2 beta + 3 gamma +version: 0 +``` + +Because the version is tracked in the transaction log, the same table supports time travel with `DeltaTable(..., version=N)` and appends that bump the version. + +## 4. Verify objects in RustFS + +List the table prefix: + +```bash +rc ls rustfs// -r +``` + +The first commit created the transaction log and one Parquet file: + +```text +events/_delta_log/00000000000000000000.json +events/part-00000-3859855e-45e4-4ae5-94ff-2d8eab5e7ebb-c000.snappy.parquet +``` + +Every new write adds a `NNNNNNNNNNNNNNNNNNNN.json` log entry and Parquet parts; readers replay the log to get a consistent snapshot. + +![Delta table files stored in the RustFS Console](./images/rustfs-delta-table.png) + +## 5. Stop or reset + +delta-rs holds no state of its own. To delete the table: + +```bash +rc rm rustfs//events/ --recursive --force +``` + +## Troubleshooting + +### `Import pyarrow failed` when writing a pandas DataFrame + +`write_deltalake` converts frames through Arrow. Install `pyarrow` alongside `deltalake` and `pandas`. + +### `Generic DeltaTable error: commit conflict` or rename errors on commit + +delta-rs commits by copying and renaming temporary objects, which requires atomic rename on the backend. For S3-compatible stores without copy-if-not-exists, set `AWS_S3_ALLOW_UNSAFE_RENAME: "true"` in `storage_options` — acceptable for a single writer, not for concurrent writers. + +### `Unknown lengthy error: AWS connectivity or endpoint errors` + +Confirm `AWS_ENDPOINT_URL` includes the scheme and that `AWS_ALLOW_HTTP` is `"true"` for plain-HTTP endpoints; without it the S3 client only speaks HTTPS. + +## Next steps + +- Review [S3 compatibility notes](/administration/protocols/s3) before adopting additional Delta clients. +- Create dedicated production credentials with [Access Key Management](/security-compliance/iam/access-token). +- Follow the [delta-rs usage documentation](https://delta-io.github.io/delta-rs/usage/writing/writing-to-s3/) for concurrent-writer setups with DynamoDB-backed commit coordination. diff --git a/content/ja/developer/integration/big-data/images/rustfs-airflow-object.png b/content/ja/developer/integration/big-data/images/rustfs-airflow-object.png new file mode 100644 index 00000000..9177d249 Binary files /dev/null and b/content/ja/developer/integration/big-data/images/rustfs-airflow-object.png differ diff --git a/content/ja/developer/integration/big-data/images/rustfs-delta-table.png b/content/ja/developer/integration/big-data/images/rustfs-delta-table.png new file mode 100644 index 00000000..093a5cf1 Binary files /dev/null and b/content/ja/developer/integration/big-data/images/rustfs-delta-table.png differ diff --git a/content/ja/developer/integration/big-data/images/rustfs-kafka-sink.png b/content/ja/developer/integration/big-data/images/rustfs-kafka-sink.png new file mode 100644 index 00000000..dd4b11b5 Binary files /dev/null and b/content/ja/developer/integration/big-data/images/rustfs-kafka-sink.png differ diff --git a/content/ja/developer/integration/big-data/index.md b/content/ja/developer/integration/big-data/index.md index d610500d..e3adb48b 100644 --- a/content/ja/developer/integration/big-data/index.md +++ b/content/ja/developer/integration/big-data/index.md @@ -8,6 +8,7 @@ Use **RustFS** as the object storage layer for data analytics systems that suppo ## Systems - [ClickHouse](./clickhouse.md) +- [Airflow](./airflow.md) - [Hudi](./hudi.md) - [Iceberg](./iceberg.md) - [PyIceberg](./pyiceberg.md) @@ -16,8 +17,10 @@ Use **RustFS** as the object storage layer for data analytics systems that suppo - [OpenDAL](./opendal.md) - [DuckDB](./duckdb.md) - [Doris](./doris.md) +- [Delta Lake](./delta-lake.md) - [lakeFS](./lakefs.md) - [InfluxDB](./influxdb.md) +- [Kafka](./kafka.md) - [Spark](./spark.md) - [Flink](./flink.md) - [Trino](./trino.md) diff --git a/content/ja/developer/integration/big-data/kafka.md b/content/ja/developer/integration/big-data/kafka.md new file mode 100644 index 00000000..ff1e92f9 --- /dev/null +++ b/content/ja/developer/integration/big-data/kafka.md @@ -0,0 +1,178 @@ +--- +title: "Kafka" +description: "Offload Kafka topic data to RustFS with the Kafka Connect S3 sink connector." +--- + +This guide connects [Apache Kafka](https://github.com/apache/kafka) — the distributed event streaming platform — to **RustFS** through the Kafka Connect S3 sink connector. You will run a KRaft broker and a Connect worker, deploy the S3 sink for a topic, and produce records that land as objects in a RustFS bucket. The workflow was verified with `apache/kafka:4.0.0` and `confluentinc/kafka-connect-s3` v10.5.25 against `rustfs/rustfs-x86-musl:v2.3.1`. + +Kafka's KIP-405 tiered storage needs a `RemoteLogStorageManager` plugin, and the S3 implementations in the ecosystem are vendor-proprietary. The Connect S3 sink is the open, self-hosted way to move topic data to S3-compatible storage and is the approach documented here. + +You need Docker. This deployment is intended for local integration testing, not production. + +## Architecture + +```mermaid +flowchart LR + Producer["Console producer"] -->|"records"| Broker["Kafka broker :9092"] + Broker -->|"consumer group"| Connect["Connect S3 sink"] + Connect -->|"batched objects"| RustFS["RustFS :9000"] +``` + +The sink task consumes a topic in a dedicated consumer group and writes record batches to the bucket, one object per `flush.size` records per partition. + +## 1. Run the broker + +Start a KRaft broker whose advertised listener is reachable from other containers: + +```bash +docker run -d --name kafka --hostname kafka --network oo-rustfs_default \ + -e CLUSTER_ID=5L6g3nShT-eMCtK--X86sw \ + -e KAFKA_NODE_ID=1 \ + -e KAFKA_PROCESS_ROLES=broker,controller \ + -e KAFKA_LISTENERS=PLAINTEXT://:9092,CONTROLLER://:9093 \ + -e KAFKA_ADVERTISED_LISTENERS=PLAINTEXT://kafka:9092 \ + -e KAFKA_CONTROLLER_LISTENER_NAMES=CONTROLLER \ + -e KAFKA_LISTENER_SECURITY_PROTOCOL_MAP=CONTROLLER:PLAINTEXT,PLAINTEXT:PLAINTEXT \ + -e KAFKA_CONTROLLER_QUORUM_VOTERS=1@kafka:9093 \ + -e KAFKA_OFFSETS_TOPIC_REPLICATION_FACTOR=1 \ + -e KAFKA_TRANSACTION_STATE_LOG_REPLICATION_FACTOR=1 \ + -e KAFKA_TRANSACTION_STATE_LOG_MIN_ISR=1 \ + apache/kafka:4.0.0 + +docker exec kafka /opt/kafka/bin/kafka-topics.sh \ + --bootstrap-server localhost:9092 \ + --create --topic rustfs-topic --partitions 1 --replication-factor 1 + +rc mb rustfs/kafka-demo +``` + +The image defaults to advertising `localhost:9092`, which only works from inside the broker container. The `KAFKA_ADVERTISED_LISTENERS` override is what makes Connect (and any remote client) able to reach the broker. + +## 2. Install the connector + +Download the Confluent Hub archive, which bundles the connector and its dependencies, and unpack it where the worker can see it: + +```bash +curl -Lo kafka-connect-s3.zip "https://hub-downloads.confluent.io/api/plugins/confluentinc/kafka-connect-s3/versions/10.5.25/confluentinc-kafka-connect-s3-10.5.25.zip" +unzip kafka-connect-s3.zip -d /opt/kafka-conn/plugins +``` + +## 3. Configure the worker and the sink + +Create the worker properties. The value converter must be `ByteArrayConverter` so records are written verbatim: + +```ini title="worker.properties" +bootstrap.servers=kafka:9092 +key.converter=org.apache.kafka.connect.storage.StringConverter +value.converter=org.apache.kafka.connect.converters.ByteArrayConverter +offset.storage.file.filename=/tmp/connect.offsets +offset.flush.interval.ms=5000 +plugin.path=/opt/kafka-conn-plugins +``` + +Create the sink connector configuration, replacing all connection placeholders: + +```ini title="rustfs-sink.properties" +name=rustfs-sink +connector.class=io.confluent.connect.s3.S3SinkConnector +tasks.max=1 +topics=rustfs-topic +s3.bucket.name=kafka-demo +s3.region=us-east-1 +store.url=http://:9000 +s3.path.style.access.enabled=true +flush.size=3 +storage.class=io.confluent.connect.s3.storage.S3Storage +format.class=io.confluent.connect.s3.format.bytearray.ByteArrayFormat +consumer.override.auto.offset.reset=earliest +``` + +Kafka 4.0 moved the class to `org.apache.kafka.connect.converters.ByteArrayConverter` — the old `storage` package path no longer resolves. + +## 4. Run the worker + +Run `connect-standalone` in the foreground so its logs go to `docker logs`, with the bucket credentials in the environment: + +```bash +docker run -d --name kafka-connect --hostname kafka-connect \ + --network oo-rustfs_default \ + -v /opt/kafka-conn/plugins:/opt/kafka-conn-plugins:ro \ + -v "$PWD/worker.properties":/etc/kafka/worker.properties:ro \ + -v "$PWD/rustfs-sink.properties":/etc/kafka/sink.properties:ro \ + -e AWS_ACCESS_KEY_ID= \ + -e AWS_SECRET_ACCESS_KEY= \ + apache/kafka:4.0.0 \ + /opt/kafka/bin/connect-standalone.sh /etc/kafka/worker.properties /etc/kafka/sink.properties +``` + +The worker is ready when the sink task claims the partition: + +```text +INFO [rustfs-sink|task-0] Assigned topic partitions: [rustfs-topic-0] +``` + +## 5. Produce records + +Send at least `flush.size` records so the connector completes a batch: + +```bash +docker exec kafka sh -c "printf 'msg-one\nmsg-two\nmsg-three\n' | \ + /opt/kafka/bin/kafka-console-producer.sh --bootstrap-server localhost:9092 --topic rustfs-topic" +``` + +After a few seconds the batch becomes an object in the bucket: + +```bash +rc ls rustfs/kafka-demo/ -r +rc cat rustfs/kafka-demo/topics/rustfs-topic/partition=0/rustfs-topic+0+0000000000.bin +``` + +```text +topics/rustfs-topic/partition=0/rustfs-topic+0+0000000000.bin +topics/rustfs-topic/partition=0/rustfs-topic+0+0000000003.bin +msg-one +msg-two +msg-three +``` + +The object name encodes topic, partition, and starting offset. Each subsequent batch of three records lands in the next object (`+0000000003.bin` and so on). + +![Kafka sink objects stored in the RustFS Console](./images/rustfs-kafka-sink.png) + +## 6. Stop or reset + +To tear down the demo while keeping the bucket objects: + +```bash +docker rm -f kafka-connect kafka +``` + +To delete the stored data: + +```bash +rc rm rustfs/kafka-demo/ --recursive --force +``` + +## Troubleshooting + +### `AdminClient ... Rebootstrapping with Cluster (id: null)` loops forever + +The broker advertises `localhost:9092`, so a remote client receives metadata pointing at itself. Set `KAFKA_ADVERTISED_LISTENERS=PLAINTEXT://kafka:9092` (and matching listener variables) as in step 1. + +### `Invalid schema type for ByteArrayConverter: STRING` + +The `ByteArrayFormat` writer only accepts raw bytes. Either switch the worker's `value.converter` to the ByteArray converter or choose a format class that matches the converter you use. + +### `Class org.apache.kafka.connect.storage.ByteArrayConverter could not be found` + +Kafka 4.0 moved the class to `org.apache.kafka.connect.converters.ByteArrayConverter`. Use the new package path in `worker.properties`. + +### The connector downloads but the plugin is not found + +The plain connector JAR from Maven lacks its dependencies. Use the Confluent Hub archive from step 2, which bundles the complete `lib/` directory, and make sure `plugin.path` points at the directory that contains the connector folder. + +## Next steps + +- Review [S3 compatibility notes](/administration/protocols/s3) before adopting additional connectors. +- Create dedicated production credentials with [Access Key Management](/security-compliance/iam/access-token). +- Follow the [Kafka Connect S3 sink documentation](https://docs.confluent.io/kafka-connect-s3/current/index.html) for Parquet and Avro formats, partitioning by time, and IAM-based credential chains. diff --git a/content/ja/developer/integration/big-data/meta.json b/content/ja/developer/integration/big-data/meta.json index 22b9c24b..ac3449c1 100644 --- a/content/ja/developer/integration/big-data/meta.json +++ b/content/ja/developer/integration/big-data/meta.json @@ -2,12 +2,15 @@ "title": "データ分析", "pages": [ "clickhouse", + "airflow", "duckdb", "doris", + "delta-lake", "flink", "hudi", "iceberg", "influxdb", + "kafka", "lakefs", "milvus", "mlflow", diff --git a/content/ja/developer/integration/devops/images/rustfs-opensearch-snapshot.png b/content/ja/developer/integration/devops/images/rustfs-opensearch-snapshot.png new file mode 100644 index 00000000..60b3bd3a Binary files /dev/null and b/content/ja/developer/integration/devops/images/rustfs-opensearch-snapshot.png differ diff --git a/content/ja/developer/integration/devops/index.md b/content/ja/developer/integration/devops/index.md index 9e31d321..612fe1a5 100644 --- a/content/ja/developer/integration/devops/index.md +++ b/content/ja/developer/integration/devops/index.md @@ -8,6 +8,7 @@ S3 互換エンドポイントをサポートする DevOps プラットフォー ## プラットフォームとツール - [Elasticsearch](./elasticsearch.md) +- [OpenSearch](./opensearch.md) - [Gitea](./gitea.md) - [Jenkins](./jenkins.md) - [Terraform](./terraform.md) diff --git a/content/ja/developer/integration/devops/meta.json b/content/ja/developer/integration/devops/meta.json index 67c84c6e..55eada6a 100644 --- a/content/ja/developer/integration/devops/meta.json +++ b/content/ja/developer/integration/devops/meta.json @@ -2,6 +2,7 @@ "title": "DevOps", "pages": [ "elasticsearch", + "opensearch", "gitea", "jenkins", "terraform" diff --git a/content/ja/developer/integration/devops/opensearch.md b/content/ja/developer/integration/devops/opensearch.md new file mode 100644 index 00000000..9d44e530 --- /dev/null +++ b/content/ja/developer/integration/devops/opensearch.md @@ -0,0 +1,172 @@ +--- +title: "OpenSearch" +description: "Snapshot OpenSearch indices to RustFS with the repository-s3 plugin." +--- + +This guide connects [OpenSearch](https://github.com/opensearch-project/OpenSearch) — the open-source search and analytics suite derived from Elasticsearch — to **RustFS** through the `repository-s3` plugin. You will register an S3 snapshot repository backed by a RustFS bucket, take a snapshot of an index, and restore it. The workflow was verified with `opensearchproject/opensearch:3.8.0` and the bundled `repository-s3` plugin against `rustfs/rustfs-x86-musl:v2.3.1`. + +You need Docker, or an OpenSearch node where you can install plugins and edit configuration. This deployment is intended for local integration testing, not production. + +## Architecture + +```mermaid +flowchart LR + Client["REST client"] --> OS["OpenSearch :9200"] + OS -->|"snapshot files"| RustFS["RustFS :9000"] + RustFS -->|"restore"| OS +``` + +The `repository-s3` plugin writes snapshots as shard archives plus metadata blobs in the bucket. Registration is cluster-wide, so every node needs the plugin and the same client configuration. + +## 1. Run OpenSearch + +Start a single node with security disabled and a small heap: + +```bash +docker run -d --name opensearch --network oo-rustfs_default -p 9200:9200 \ + -e discovery.type=single-node \ + -e OPENSEARCH_JAVA_OPTS="-Xms512m -Xmx512m" \ + -e DISABLE_SECURITY_PLUGIN=true \ + opensearchproject/opensearch:3.8.0 +``` + +The node is ready when `curl http://localhost:9200` returns the cluster header (allow one to two minutes). + +## 2. Install the repository-s3 plugin + +The S3 repository plugin is not preloaded. Install it and restart the node: + +```bash +docker exec opensearch bin/opensearch-plugin install --batch repository-s3 +docker restart opensearch +``` + +## 3. Configure the S3 client + +Credentials are secure settings: they belong in the OpenSearch keystore, not in the repository request or `opensearch.yml`. Create the keystore entries, replacing all connection placeholders: + +```bash +docker exec opensearch sh -c \ + "printf '' | bin/opensearch-keystore create 2>/dev/null; \ + printf '' | bin/opensearch-keystore add -f -x s3.client.default.access_key; \ + printf '' | bin/opensearch-keystore add -f -x s3.client.default.secret_key" +``` + +Add the non-secure client settings to `config/opensearch.yml`: + +```yaml title="opensearch.yml" +network.host: 0.0.0.0 +plugins.security.disabled: true +s3.client.default.endpoint: http://:9000 +s3.client.default.protocol: http +s3.client.default.path_style_access: "true" +``` + +Restart the node once more so it reads both the keystore and the new settings: + +```bash +docker restart opensearch +``` + +Create the bucket while the node boots: + +```bash +rc mb rustfs/opensearch-snapshots +``` + +## 4. Register the repository and snapshot + +Create a test index with a document, then register the repository: + +```bash +curl -sX PUT http://localhost:9200/rustfs-demo -H "Content-Type: application/json" \ + -d '{"settings":{"number_of_shards":1}}' + +curl -sX PUT http://localhost:9200/rustfs-demo/_doc/1 -H "Content-Type: application/json" \ + -d '{"product":"rustfs","via":"opensearch-snapshot"}' + +curl -sX PUT "http://localhost:9200/_snapshot/rustfs-repo" -H "Content-Type: application/json" \ + -d '{"type":"s3","settings":{"bucket":"opensearch-snapshots","region":"us-east-1","server_side_encryption_type":"bucket_default"}}' +``` + +The `server_side_encryption_type: bucket_default` setting matters: without it the plugin requests SSE-S3, which a self-hosted RustFS without a server-side encryption master key rejects. + +Take a snapshot and wait for completion: + +```bash +curl -sX PUT "http://localhost:9200/_snapshot/rustfs-repo/snapshot-1?wait_for_completion=true" \ + -H "Content-Type: application/json" -d '{"indices":"rustfs-demo"}' +``` + +```text +{"snapshot":{"snapshot":"snapshot-1","state":"SUCCESS","indices":["rustfs-demo"],...}} +``` + +## 5. Verify objects and restore + +List the bucket: + +```bash +rc ls rustfs/opensearch-snapshots/ -r +``` + +```text +index-0 +index.latest +indices/5x1bwsWaSv2XINIbeoe-RQ/0/__GgxvoCw-TKuMMAYBq5Khag +indices/5x1bwsWaSv2XINIbeoe-RQ/0/snap-kRgFBiuyQIyPo3_-CMp7Hw.dat +meta-kRgFBiuyQIyPo3_-CMp7Hw.dat +snap-kRgFBiuyQIyPo3_-CMp7Hw.dat +``` + +Delete the index and restore it from the snapshot: + +```bash +curl -sX DELETE http://localhost:9200/rustfs-demo +curl -sX POST "http://localhost:9200/_snapshot/rustfs-repo/snapshot-1/_restore?wait_for_completion=true" +curl -s http://localhost:9200/rustfs-demo/_doc/1 +``` + +```text +{"_index":"rustfs-demo","_id":"1","found":true,"_source":{"product":"rustfs","via":"opensearch-snapshot"}} +``` + +![OpenSearch snapshot stored in the RustFS Console](./images/rustfs-opensearch-snapshot.png) + +## 6. Stop or reset + +To tear down the demo while keeping the bucket objects: + +```bash +docker rm -f opensearch +``` + +To delete the stored snapshots: + +```bash +rc rm rustfs/opensearch-snapshots/ --recursive --force +``` + +## Troubleshooting + +### `Setting [access_key] is insecure, but property [allow_insecure_settings] is not set` + +Inline credentials in the repository request are rejected. Store them in the keystore as shown in step 3 — `access_key` and `secret_key` are secure settings in OpenSearch. + +### `SSE-S3 requires RUSTFS_SSE_S3_MASTER_KEY ... (Status Code: 400)` + +The plugin encrypts uploads with SSE-S3 by default. Register the repository with `"server_side_encryption_type": "bucket_default"` so no encryption header is sent, as in step 4. + +### `unknown setting [s3.client.default.access_key]` at startup + +The settings reference the repository-s3 plugin. If the node fails to start with them present, the plugin is not installed in that container — repeat step 2 (a fresh container loses plugins installed with `docker exec`). + +### Repository verification fails with `path is not accessible` + +The node cannot reach the bucket: check that `s3.client.default.endpoint` is reachable from the container, `path_style_access` is `"true"`, and the keystore credentials were loaded (they are read at startup — restart after adding them). + +## Next steps + +- Review [S3 compatibility notes](/administration/protocols/s3) before adopting additional OpenSearch repositories. +- Create dedicated production credentials with [Access Key Management](/security-compliance/iam/access-token). +- Follow the [OpenSearch snapshots documentation](https://docs.opensearch.org/docs/latest/tuning-your-cluster/availability-and-recovery/snapshots/index/) to automate snapshots with Snapshot Management (SM) policies. diff --git a/content/ja/developer/integration/index.md b/content/ja/developer/integration/index.md index 1746d641..54d85ff5 100644 --- a/content/ja/developer/integration/index.md +++ b/content/ja/developer/integration/index.md @@ -7,14 +7,14 @@ description: "RustFS をリバースプロキシ、バックアップツール ## Integration categories -- [Reverse Proxy](./reverse-proxy/index.md) は Nginx、Traefik、Caddy、HAProxy を扱います。 +- [Reverse Proxy](./reverse-proxy/index.md) は Nginx、Traefik、Caddy、HAProxy、Envoy を扱います。 - [Backup](./backup/index.md) は Kopia、Longhorn、Restic、Velero を扱います。 -- [AI](./ai/index.md) は Ray などの AI プラットフォームを扱います。 -- [データ分析](./big-data/index.md) は ClickHouse、Doris、Hudi、Iceberg、lakeFS、Milvus、OpenDAL、Vitess、Zeppelin などの分析システムを扱います。 +- [AI](./ai/index.md) は Ray、vLLM などの AI プラットフォームを扱います。 +- [データ分析](./big-data/index.md) は Airflow、ClickHouse、Delta Lake、Doris、Hudi、Iceberg、Kafka、lakeFS、Milvus、OpenDAL、Vitess、Zeppelin などの分析システムを扱います。 - [クラウドネイティブ](./cloud-native/index.md) は Cortex、Flux を扱います。 -- [オブザーバビリティ](./observability/index.md) は Fluentd、OpenObserve、OpenTelemetry、Thanos、Tempo などのテレメトリシステムを扱います。 -- [その他](./others/index.md) はコミュニティ主導の Python 用 capo SDK を扱います。 +- [オブザーバビリティ](./observability/index.md) は Fluentd、GreptimeDB、Loki、OpenObserve、OpenTelemetry、Tempo、Thanos、VictoriaMetrics などのテレメトリシステムを扱います。 +- [その他](./others/index.md) は capo SDK、rclone、JuiceFS、Nextcloud、tusd などのツールを扱います。 - [コンテナレジストリ](./registry/index.md) は Harbor を扱います。 -- [DevOps](./devops/index.md) は Elasticsearch、Gitea、Jenkins、Terraform を扱います。 +- [DevOps](./devops/index.md) は Elasticsearch、Gitea、Jenkins、OpenSearch、Terraform を扱います。 各ガイドでは、連携先システムを設定する際に使用する RustFS のエンドポイントとアドレス指定の要件を示します。 \ No newline at end of file diff --git a/content/ja/developer/integration/observability/greptimedb.md b/content/ja/developer/integration/observability/greptimedb.md new file mode 100644 index 00000000..5b7ee690 --- /dev/null +++ b/content/ja/developer/integration/observability/greptimedb.md @@ -0,0 +1,129 @@ +--- +title: "GreptimeDB" +description: "Run GreptimeDB with RustFS as the S3-compatible object storage backend." +--- + +This guide connects [GreptimeDB](https://github.com/GreptimeTeam/greptimedb) — the open-source, cloud-native time-series database — to **RustFS** as its object storage backend. You will start a standalone instance with its `[storage]` section pointed at a RustFS bucket, write time-series rows through the SQL API, and confirm the Parquet files and manifests in the bucket. The workflow was verified with `greptime/greptimedb` (main, commit `179ff8e5`) against `rustfs/rustfs-x86-musl:v2.3.1`. + +You need Docker, or a local GreptimeDB binary. This deployment is intended for local integration testing, not production. + +## Architecture + +```mermaid +flowchart LR + SQL["SQL / Prometheus API"] --> DB["GreptimeDB"] + DB -->|"SST + manifests"| RustFS["RustFS :9000"] +``` + +GreptimeDB keeps its write-ahead log and recent data locally, then persists SSTables (Parquet) and table manifests to object storage. Pointing the storage backend at RustFS makes the bucket the durable home of all table data. + +## 1. Configure the storage backend + +Create the bucket and a config file with an S3 storage section, replacing all connection placeholders: + +```toml title="greptimedb.toml" +[storage] +type = "S3" +bucket = "" +root = "greptimedb" +access_key_id = "" +secret_access_key = "" +endpoint = "http://:9000" +region = "us-east-1" +``` + +GreptimeDB uses path-style requests for custom endpoints by default; virtual-hosted style must be opted into explicitly with `enable_virtual_host_style`, so no extra flag is needed for RustFS. + +## 2. Run GreptimeDB + +Start a standalone instance with the config file: + +```bash +docker run -d --name greptimedb --network oo-rustfs_default -p 4000:4000 -p 4002:4002 \ + -v "$PWD/greptimedb.toml":/etc/greptimedb/greptimedb.toml:ro \ + greptime/greptimedb:latest standalone start \ + --http-addr 0.0.0.0:4000 \ + --mysql-addr 0.0.0.0:4002 \ + --config-file /etc/greptimedb/greptimedb.toml +``` + +Port `4000` serves the HTTP SQL endpoint and `4002` the MySQL protocol. + +## 3. Write and query time series + +Create a table, insert rows, and read them back. The HTTP SQL endpoint takes form-encoded requests: + +```bash +curl -s -X POST "http://localhost:4000/v1/sql" \ + --data-urlencode "sql=CREATE TABLE rustfs_demo (host STRING, cpu DOUBLE, mem DOUBLE, ts TIMESTAMP TIME INDEX)" + +curl -s -X POST "http://localhost:4000/v1/sql" \ + --data-urlencode "sql=INSERT INTO rustfs_demo VALUES (\"node-1\", 0.31, 0.62, 1790681000000), (\"node-1\", 0.35, 0.63, 1790681060000), (\"node-2\", 0.51, 0.71, 1790681000000)" +``` + +```text +{"output":[{"affectedrows":3}],"execution_time_ms":2} +``` + +Query the rows back: + +```bash +curl -s -X POST "http://localhost:4000/v1/sql" \ + --data-urlencode "sql=SELECT * FROM rustfs_demo ORDER BY ts" +``` + +```text +{"output":[{"records":{"rows":[["node-2",0.51,0.71,1790681000000],["node-1",0.35,0.63,1790681060000]],"total_rows":2}}]} +``` + +## 4. Verify objects in RustFS + +List the bucket — after the memtable flushes, the bucket holds Parquet SSTables and JSON manifests: + +```bash +rc ls rustfs// -r +``` + +```text +greptimedb/data/greptime/public/1024/1024_0000000000/manifest/00000000000000000000.json +greptimedb/data/greptime/greptime_private/1025/1025_0000000000/b11e8b25-5763-4f05-bcab-b6ee0a756a69.parquet +greptimedb/data/greptime/greptime_private/1025/1025_0000000000/manifest/00000000000000000001.json +``` + +Each database gets a directory under `data/`, and per-region `manifest/*.json` files describe the SSTables GreptimeDB reads back during queries. + +![GreptimeDB data stored in the RustFS Console](./images/rustfs-greptimedb-data.png) + +## 5. Stop or reset + +To tear down the demo while keeping the bucket objects: + +```bash +docker rm -f greptimedb +``` + +To delete the stored data: + +```bash +rc rm rustfs// --recursive --force +``` + +## Troubleshooting + +### `Form requests must have Content-Type: application/x-www-form-urlencoded` + +The `/v1/sql` HTTP endpoint only accepts form-encoded bodies. Pass SQL with `curl --data-urlencode "sql=..."` (or `application/x-www-form-urlencoded`), not as a JSON body. + +### Bucket stays empty + +GreptimeDB flushes memtables to object storage asynchronously. Run a few more inserts and wait a few seconds, or trigger a manual flush, then list the bucket again. + +### Startup fails with an S3 error + +Confirm `endpoint` includes the scheme, the bucket exists, and `access_key_id`/`secret_access_key` match a RustFS access key. The `root` value is optional but keeps the table tree under a known prefix. + +## Next steps + +- Review [S3 compatibility notes](/administration/protocols/s3) before adopting additional GreptimeDB storage options. +- Create dedicated production credentials with [Access Key Management](/security-compliance/iam/access-token). +- Follow the [GreptimeDB configuration reference](https://docs.greptime.com/operational-guide/configure/configure-datanode/) to tune flush intervals and cache layers for production workloads. diff --git a/content/ja/developer/integration/observability/images/rustfs-greptimedb-data.png b/content/ja/developer/integration/observability/images/rustfs-greptimedb-data.png new file mode 100644 index 00000000..1874f210 Binary files /dev/null and b/content/ja/developer/integration/observability/images/rustfs-greptimedb-data.png differ diff --git a/content/ja/developer/integration/observability/images/rustfs-vm-backups.png b/content/ja/developer/integration/observability/images/rustfs-vm-backups.png new file mode 100644 index 00000000..a0bd0cee Binary files /dev/null and b/content/ja/developer/integration/observability/images/rustfs-vm-backups.png differ diff --git a/content/ja/developer/integration/observability/index.md b/content/ja/developer/integration/observability/index.md index 56acb094..ae1e264d 100644 --- a/content/ja/developer/integration/observability/index.md +++ b/content/ja/developer/integration/observability/index.md @@ -8,10 +8,12 @@ S3 互換エンドポイントをサポートするオブザーバビリティ ## プラットフォーム - [Fluentd](./fluentd.md) +- [GreptimeDB](./greptimedb.md) - [OpenObserve](./openobserve.md) - [OpenTelemetry](./opentelemetry.md) - [Loki](./loki.md) - [Tempo](./tempo.md) - [Thanos](./thanos.md) +- [VictoriaMetrics](./victoriametrics.md) テレメトリデータは専用バケットに保存し、必要なバケット操作のみに権限が絞られた認証情報を使用してください。 diff --git a/content/ja/developer/integration/observability/meta.json b/content/ja/developer/integration/observability/meta.json index 1f09554f..c878b558 100644 --- a/content/ja/developer/integration/observability/meta.json +++ b/content/ja/developer/integration/observability/meta.json @@ -2,10 +2,12 @@ "title": "オブザーバビリティ", "pages": [ "fluentd", + "greptimedb", "loki", "openobserve", "opentelemetry", "tempo", - "thanos" + "thanos", + "victoriametrics" ] } diff --git a/content/ja/developer/integration/observability/victoriametrics.md b/content/ja/developer/integration/observability/victoriametrics.md new file mode 100644 index 00000000..efc799d9 --- /dev/null +++ b/content/ja/developer/integration/observability/victoriametrics.md @@ -0,0 +1,157 @@ +--- +title: "VictoriaMetrics" +description: "Back up VictoriaMetrics snapshots to RustFS with vmbackup." +--- + +This guide connects [VictoriaMetrics](https://github.com/VictoriaMetrics/VictoriaMetrics) — the Prometheus-compatible time-series database — to **RustFS** through `vmbackup` and `vmrestore`. You will run a single-node instance, import metrics, create an instant snapshot, back it up to a RustFS bucket, and restore the data into a fresh directory. The workflow was verified with `victoria-metrics`, `vmbackup`, and `vmrestore` v1.x images against `rustfs/rustfs-x86-musl:v2.3.1`. + +You need Docker. This deployment is intended for local integration testing, not production. + +## Architecture + +```mermaid +flowchart LR + Import["Prometheus import API"] --> VM["VictoriaMetrics :8428"] + VM -->|"instant snapshot"| Backup["vmbackup"] + Backup -->|"S3 API"| RustFS["RustFS :9000"] + RustFS -->|"restore"| Restore["vmrestore"] +``` + +`vmbackup` uploads a consistent point-in-time snapshot of the storage directory to any S3-compatible endpoint. `vmrestore` reverses the process, producing a data directory a VictoriaMetrics instance can open directly. + +## 1. Run VictoriaMetrics + +Create the bucket and start a single-node instance: + +```bash +rc mb rustfs/vm-backups + +docker run -d --name vm --network oo-rustfs_default -p 8428:8428 \ + -v vm-data:/storage \ + victoriametrics/victoria-metrics:latest \ + -storageDataPath=/storage -retentionPeriod=100y +``` + +## 2. Import metrics + +Write a couple of samples through the Prometheus import API: + +```bash +echo "vm_demo_metric 123" | curl -s --data-binary @- http://localhost:8428/api/v1/import/prometheus +echo "vm_demo_metric 456" | curl -s --data-binary @- http://localhost:8428/api/v1/import/prometheus +``` + +The endpoint answers `204 No Content`. Confirm the data is queryable: + +```bash +curl -s "http://localhost:8428/api/v1/export?match[]=vm_demo_metric" +``` + +```text +{"metric":{"__name__":"vm_demo_metric"},"values":[123,456],"timestamps":[1790680894604,1790680894619]} +``` + +## 3. Create a snapshot + +Ask VictoriaMetrics for a consistent snapshot: + +```bash +curl -s http://localhost:8428/snapshot/create +``` + +```text +{"status":"ok","snapshot":"20260929112134-18D9C6C76704F913"} +``` + +## 4. Back the snapshot up to RustFS + +Run `vmbackup` against the same storage volume, replacing the credential placeholders. The snapshot name comes from step 3: + +```bash +docker run --rm --network oo-rustfs_default \ + -e AWS_ACCESS_KEY_ID= \ + -e AWS_SECRET_ACCESS_KEY= \ + --volumes-from vm \ + victoriametrics/vmbackup:latest \ + -storageDataPath=/storage \ + -snapshotName=20260929112134-18D9C6C76704F913 \ + -dst=s3://vm-backups/demo \ + -customS3Endpoint=http://:9000 +``` + +```text +backup ... to S3{bucket: "vm-backups", dir: "demo/"} is complete; uploaded 760 bytes +``` + +`-customS3Endpoint` redirects the AWS SDK to RustFS; custom endpoints are addressed with path-style requests automatically. Set `AWS_EC2_METADATA_DISABLED=true` on hosts without an EC2 metadata service to skip credential lookup delays. + +## 5. Verify and restore + +List the bucket prefix: + +```bash +rc ls rustfs/vm-backups/demo/ +``` + +```text +backup_complete.ignore +backup_metadata.ignore +data/ +metadata/ +``` + +`backup_complete.ignore` marks a complete backup. Restore it into a fresh directory: + +```bash +docker run --rm --network oo-rustfs_default \ + -e AWS_ACCESS_KEY_ID= \ + -e AWS_SECRET_ACCESS_KEY= \ + -v /opt/vm-restore:/restore \ + victoriametrics/vmrestore:latest \ + -src=s3://vm-backups/demo \ + -storageDataPath=/restore \ + -customS3Endpoint=http://:9000 +``` + +```text +restored 760 bytes from backup in 0.055 seconds +``` + +The restored directory contains `data/`, `metadata/`, and a lock file — exactly what a VictoriaMetrics instance expects at `-storageDataPath`. + +![VictoriaMetrics backup stored in the RustFS Console](./images/rustfs-vm-backups.png) + +## 6. Stop or reset + +To tear down the demo while keeping the bucket objects: + +```bash +docker rm -f vm +docker volume rm vm-data +``` + +To delete the stored backups: + +```bash +rc rm rustfs/vm-backups/ --recursive --force +``` + +## Troubleshooting + +### `vmbackup` hangs at startup or fails to find credentials + +The AWS SDK probes the EC2 metadata service when environment credentials are absent. On machines without IMDS, export `AWS_EC2_METADATA_DISABLED=true` next to the key variables. + +### Backup parts re-upload on every run + +`vmbackup` performs incremental backups by comparing local and remote file hashes. Restoring to a fresh directory and running `vmbackup` from there re-uploads everything; keep the original data directory for incremental runs. + +### Query returns nothing right after import + +Imports are accepted asynchronously and the instant query endpoint can lag on a busy single node. Verify with `/api/v1/export` (or wait a few seconds) before creating the snapshot. + +## Next steps + +- Review [S3 compatibility notes](/administration/protocols/s3) before adopting additional VictoriaMetrics components. +- Create dedicated production credentials with [Access Key Management](/security-compliance/iam/access-token). +- Follow the [vmbackup documentation](https://docs.victoriametrics.com/vmbackup/) to schedule backups and prune old snapshots. diff --git a/content/ja/developer/integration/others/images/rustfs-juicefs-chunks.png b/content/ja/developer/integration/others/images/rustfs-juicefs-chunks.png new file mode 100644 index 00000000..737490b6 Binary files /dev/null and b/content/ja/developer/integration/others/images/rustfs-juicefs-chunks.png differ diff --git a/content/ja/developer/integration/others/images/rustfs-nextcloud-file.png b/content/ja/developer/integration/others/images/rustfs-nextcloud-file.png new file mode 100644 index 00000000..506e0d79 Binary files /dev/null and b/content/ja/developer/integration/others/images/rustfs-nextcloud-file.png differ diff --git a/content/ja/developer/integration/others/images/rustfs-rclone-sync.png b/content/ja/developer/integration/others/images/rustfs-rclone-sync.png new file mode 100644 index 00000000..dcf13712 Binary files /dev/null and b/content/ja/developer/integration/others/images/rustfs-rclone-sync.png differ diff --git a/content/ja/developer/integration/others/images/rustfs-tus-uploads.png b/content/ja/developer/integration/others/images/rustfs-tus-uploads.png new file mode 100644 index 00000000..e82a55e1 Binary files /dev/null and b/content/ja/developer/integration/others/images/rustfs-tus-uploads.png differ diff --git a/content/ja/developer/integration/others/index.md b/content/ja/developer/integration/others/index.md index eb12c5ee..a5d309cf 100644 --- a/content/ja/developer/integration/others/index.md +++ b/content/ja/developer/integration/others/index.md @@ -8,3 +8,7 @@ description: "その他の RustFS 統合。現在はコミュニティ主導の ## ガイド - [capo (Python)](./capo.md) — コミュニティ主導の capo SDK を同期・非同期クライアントで RustFS に接続します。 +rclone](./rclone.md) — sync, mount, and serve RustFS buckets from the command line. +- [tusd](./tusd.md) — receive resumable uploads into a RustFS bucket over the tus protocol. +- [JuiceFS](./juicefs.md) — mount a POSIX filesystem backed by a RustFS bucket. +- [Nextcloud](./nextcloud.md) — use RustFS as S3 external storage for Nextcloud files. diff --git a/content/ja/developer/integration/others/juicefs.md b/content/ja/developer/integration/others/juicefs.md new file mode 100644 index 00000000..428b08ce --- /dev/null +++ b/content/ja/developer/integration/others/juicefs.md @@ -0,0 +1,132 @@ +--- +title: "JuiceFS" +description: "Build a POSIX filesystem on RustFS with JuiceFS S3 object storage." +--- + +This guide connects [JuiceFS](https://github.com/juicedata/juicefs) — the cloud-native distributed POSIX filesystem — to **RustFS** as its object storage backend. You will format a volume whose data chunks live in a RustFS bucket, mount it locally, and read and write files through the mount. The workflow was verified with `juicefs v1.3.1` (community edition, SQLite metadata engine) against `rustfs/rustfs-x86-musl:v2.3.1`. + +You need the JuiceFS binary, a metadata engine, and FUSE (`fuse3` on Linux). This deployment is intended for local integration testing, not production. + +## Architecture + +```mermaid +flowchart LR + Mount["/mnt/jfs"] -->|"POSIX"| JuiceFS["JuiceFS client"] + JuiceFS -->|"metadata"| Meta["SQLite / Redis"] + JuiceFS -->|"data chunks"| RustFS["RustFS :9000"] +``` + +JuiceFS splits every file into chunks and stores them as objects under `chunks/` in the bucket, while the metadata engine tracks names, inodes, and layout. The filesystem behaves like a local disk but holds no data locally. + +## 1. Format the volume + +Create the bucket and format a JuiceFS volume backed by RustFS, replacing all connection placeholders. The bucket URL carries the endpoint, which selects path-style addressing: + +```bash +rc mb rustfs/jfs-demo + +juicefs format \ + --storage s3 \ + --bucket http://:9000/jfs-demo \ + --access-key \ + --secret-key \ + sqlite3:///opt/juicefs/jfs.db \ + rustfs-jfs +``` + +```text +Data use s3://:9000/jfs-demo/rustfs-jfs/ + OK, rustfs-jfs is ready +``` + +`sqlite3:///opt/juicefs/jfs.db` is the metadata engine for this test. In production, use Redis, MySQL, or PostgreSQL instead so multiple clients can mount the same volume. + +## 2. Mount the volume + +Mount the filesystem with the same metadata URL: + +```bash +mkdir -p /mnt/jfs +juicefs mount -d sqlite3:///opt/juicefs/jfs.db /mnt/jfs +``` + +```text +OK, rustfs-jfs is ready at /mnt/jfs +``` + +The `-d` flag runs the mount in the background. The volume is now a POSIX filesystem. + +## 3. Read and write files + +Use the mount like any other directory: + +```bash +echo "hello rustfs jfs" > /mnt/jfs/hello.txt +dd if=/dev/urandom of=/mnt/jfs/blob.bin bs=1M count=3 +mkdir -p /mnt/jfs/dir1 && echo nested > /mnt/jfs/dir1/nested.txt +cat /mnt/jfs/hello.txt +``` + +```text +hello rustfs jfs +``` + +Inspect the volume with `juicefs info`: + +```text +/mnt/jfs : + inode: 1 + files: 2 + dirs: 2 + length: 3.00 MiB +``` + +## 4. Verify chunks in RustFS + +List the bucket prefixes: + +```bash +rc ls rustfs/jfs-demo/ -r +``` + +Every file was split into content-addressed chunk objects: + +```text +rustfs-jfs/chunks/0/0/1_0_17 +rustfs-jfs/chunks/0/0/3_0_3145728 +rustfs-jfs/chunks/0/0/4_0_7 +``` + +The chunk name encodes the inode, chunk index, and size — for example `3_0_3145728` is the 3 MiB file written in step 3. + +![JuiceFS data chunks stored in the RustFS Console](./images/rustfs-juicefs-chunks.png) + +## 5. Stop or reset + +Unmount the volume, then optionally wipe the volume metadata and bucket data: + +```bash +juicefs umount /mnt/jfs +juicefs destroy --force sqlite3:///opt/juicefs/jfs.db rustfs-jfs +rc rm rustfs/jfs-demo/ --recursive --force +``` + +## Troubleshooting + +### `unknown option: --daemon` + +The background flag is a single dash: `juicefs mount -d`. Running without it keeps the mount in the foreground (useful for debugging). + +### `fusermount3: mount failed: Permission denied` + +Mounting requires the FUSE device. Inside a container, add `--device /dev/fuse --cap-add SYS_ADMIN` (or `--privileged`); on a host, install `fuse3` and confirm `/dev/fuse` exists. + +### Mount hangs or fails to reach storage + +The client must reach both the metadata engine and the bucket endpoint. Because the bucket URL embeds the endpoint, verify it from the mounting host with `curl` before formatting. + +## Next steps + +- Review [S3 compatibility notes](/administration/protocols/s3) before adopting additional JuiceFS storage backends. +- Create dedicated production credentials with [Access Key Management](/security-compliance/iam/access-token). +- Follow the [JuiceFS documentation](https://juicefs.com/docs/community/quick_start_guide/) to switch the metadata engine to Redis and mount the volume from multiple clients. diff --git a/content/ja/developer/integration/others/meta.json b/content/ja/developer/integration/others/meta.json index eb379803..186c1ffc 100644 --- a/content/ja/developer/integration/others/meta.json +++ b/content/ja/developer/integration/others/meta.json @@ -1,6 +1,10 @@ { "title": "その他", "pages": [ - "capo" + "capo", + "juicefs", + "nextcloud", + "rclone", + "tusd" ] } diff --git a/content/ja/developer/integration/others/nextcloud.md b/content/ja/developer/integration/others/nextcloud.md new file mode 100644 index 00000000..f506d75c --- /dev/null +++ b/content/ja/developer/integration/others/nextcloud.md @@ -0,0 +1,145 @@ +--- +title: "Nextcloud" +description: "Use RustFS as S3 external storage for Nextcloud files." +--- + +This guide connects [Nextcloud](https://github.com/nextcloud/server) — the self-hosted content collaboration platform — to **RustFS** through its External Storage app with the S3 backend. You will enable `files_external`, mount a RustFS bucket into every user's files view, and upload a file through WebDAV that lands directly in the bucket. The workflow was verified with `nextcloud:32.0.15` (SQLite, single container) against `rustfs/rustfs-x86-musl:v2.3.1`. + +You need Docker, or an existing Nextcloud instance with `occ` access. This deployment is intended for local integration testing, not production. + +## Architecture + +```mermaid +flowchart LR + User["Browser / WebDAV"] --> Nextcloud["Nextcloud"] + Nextcloud -->|"files_external (S3)"| RustFS["RustFS :9000"] +``` + +Nextcloud proxies file operations on the mount point to the S3 backend. Objects are stored under their mount-relative paths, so the bucket mirrors the names users see. + +## 1. Install Nextcloud + +Run Nextcloud with an admin account, replacing all connection placeholders. SQLite keeps the test self-contained; use MariaDB or PostgreSQL in production: + +```bash +docker run -d --name nextcloud --network oo-rustfs_default -p 8080:80 \ + -e NEXTCLOUD_ADMIN_USER= \ + -e NEXTCLOUD_ADMIN_PASSWORD= \ + nextcloud:32.0.15 +``` + +If the web UI still shows the installer after startup, finish it manually: + +```bash +docker exec -u www-data nextcloud php occ maintenance:install \ + --admin-user --admin-password +``` + +## 2. Enable the External Storage app + +The `files_external` app ships with Nextcloud but starts disabled, and its `occ` commands only exist once the app is enabled: + +```bash +docker exec -u www-data nextcloud php occ app:enable files_external +``` + +```text +files_external 1.24.1 enabled +``` + +## 3. Mount the RustFS bucket + +Create an external storage of backend type `amazons3` with the `amazons3::accesskey` authentication backend. Replace all connection placeholders: + +```bash +docker exec -u www-data nextcloud php occ files_external:create \ + /rustfs amazons3 amazons3::accesskey \ + --user \ + --config bucket= \ + --config hostname= \ + --config port=9000 \ + --config use_ssl=false \ + --config use_path_style=true \ + --config key= \ + --config secret= +``` + +```text +Storage created with id 1 +``` + +The mount point `/rustfs` appears in the files view of the given user. `use_path_style=true` is required for a non-AWS endpoint. Check the connection before using it: + +```bash +docker exec -u www-data nextcloud php occ files_external:verify 1 +``` + +```text + - status: ok + - code: 0 +``` + +## 4. Upload a file and verify + +Upload through the WebDAV endpoint, which writes through the external storage: + +```bash +echo "nextcloud writes to rustfs" > /tmp/nc-demo.txt + +curl -u : \ + -T /tmp/nc-demo.txt \ + http://localhost:8080/remote.php/dav/files//rustfs/nc-demo.txt \ + -o /dev/null -w "%{http_code}\n" +``` + +```text +201 +``` + +Read it back through the same path, then confirm the object in RustFS: + +```bash +rc ls rustfs// -r +``` + +```text +[2026-09-29 13:42:48] 27 B nc-demo.txt +``` + +The object key equals the path inside the mount, so files uploaded through Nextcloud can also be read directly with any S3 client. + +![Nextcloud file stored in the RustFS Console](./images/rustfs-nextcloud-file.png) + +## 5. Stop or reset + +To remove the mount without touching the bucket: + +```bash +docker exec -u www-data nextcloud php occ files_external:delete 1 +``` + +To delete the bucket contents: + +```bash +rc rm rustfs// --recursive --force +``` + +## Troubleshooting + +### `There are no commands defined in the "files_external" namespace` + +The app is not enabled yet. Run `occ app:enable files_external` first; the `occ files_external:*` commands only register afterwards. + +### `Not enough arguments (missing: "authentication_backend")` + +`files_external:create` takes the storage backend and the authentication backend as two separate arguments: `amazons3 amazons3::accesskey`. The backend identifiers are listed by `occ files_external:backends`. + +### Mount shows but is empty, or uploads fail + +Confirm `hostname` is reachable from the Nextcloud container (use the container network name, not `localhost`), `use_path_style` is `true`, and the bucket exists. `occ files_external:verify ` reports the exact connection error. + +## Next steps + +- Review [S3 compatibility notes](/administration/protocols/s3) before adopting additional external storage backends. +- Create dedicated production credentials with [Access Key Management](/security-compliance/iam/access-token). +- Follow the [Nextcloud external storage documentation](https://docs.nextcloud.com/server/latest/admin_manual/configuration_files/external_storage_configuration_gui.html) to share the mount with groups and enable versioning. diff --git a/content/ja/developer/integration/others/rclone.md b/content/ja/developer/integration/others/rclone.md new file mode 100644 index 00000000..34f9fe4a --- /dev/null +++ b/content/ja/developer/integration/others/rclone.md @@ -0,0 +1,144 @@ +--- +title: "rclone" +description: "Sync, mount, and serve RustFS buckets with rclone over its S3-compatible API." +--- + +This guide connects [rclone](https://github.com/rclone/rclone) — the command-line tool for syncing files to and from cloud storage — to **RustFS** through its S3 backend. You will configure an S3 remote for RustFS, copy and sync files, read objects back, publish a bucket over HTTP with `rclone serve`, and mount the bucket as a local filesystem with `rclone mount`. The workflow was verified with `rclone v1.75.1` against `rustfs/rustfs-x86-musl:v2.3.1`. + +You need Docker, or a local rclone binary. This deployment is intended for local integration testing, not production. + +## Architecture + +```mermaid +flowchart LR + Files["Local files"] -->|"copy / sync"| Remote["rclone S3 remote"] + Remote -->|"S3 API"| RustFS["RustFS :9000"] + RustFS -->|"mount / serve"| Client["FUSE mount / HTTP clients"] +``` + +One remote definition drives every rclone command: data transfer, mounting, and serving all use the same S3 connection. + +## 1. Configure the remote + +Create an rclone config file with an S3 remote for RustFS, replacing all connection placeholders. The `Other` provider disables AWS-specific behavior, and path-style addressing is used automatically for custom endpoints: + +```ini title="rclone.conf" +[rustfs] +type = s3 +provider = Other +access_key_id = +secret_access_key = +endpoint = http://:9000 +region = us-east-1 +``` + +## 2. Copy and read objects + +Create the bucket and upload a directory with `rclone copy`: + +```bash +rc mb rustfs/rclone-demo +rclone copy /data rustfs:rclone-demo/seed +``` + +List and read back: + +```bash +rclone ls rustfs:rclone-demo/seed +rclone cat rustfs:rclone-demo/seed/hello.txt +``` + +```text + 3145728 blob.bin + 18 hello.txt +hello from rclone +``` + +`rclone lsd rustfs:` lists every bucket on the endpoint. + +## 3. Sync a directory + +`rclone sync` makes the destination identical to the source, including deletions. Remove a local file and sync: + +```bash +rm /data/hello.txt +rclone sync /data rustfs:rclone-demo/seed +rclone lsf rustfs:rclone-demo/seed +``` + +```text +blob.bin +``` + +`hello.txt` disappears from the bucket. Add `--dry-run` first to preview the changes without touching the bucket. + +## 4. Serve a bucket over HTTP + +Publish the bucket contents as an HTTP file server: + +```bash +rclone serve http --addr 0.0.0.0:8080 rustfs:rclone-demo/seed +``` + +Any HTTP client can now download objects: + +```bash +curl -s http://localhost:8080/blob.bin -o /dev/null -w "%{http_code} %{size_download} bytes\n" +``` + +```text +200 3145728 bytes +``` + +`rclone serve` also supports WebDAV, SFTP, and S3 endpoints over the same remote. + +## 5. Mount the bucket as a filesystem + +With FUSE available, mount the bucket locally and use it like a directory: + +```bash +rclone mount rustfs:rclone-demo /mnt/rclone --daemon +ls /mnt/rclone/seed +echo test > /mnt/rclone/write-test.txt +cat /mnt/rclone/write-test.txt +``` + +Files written through the mount appear in RustFS as regular objects: + +```bash +rc ls rustfs/rclone-demo/ -r +``` + +```text +[2026-09-29 11:12:41] 5 B write-test.txt +``` + +Unmount with `fusermount -u /mnt/rclone` when finished. + +## 6. Stop or reset + +rclone holds no server-side state. To delete the demo data: + +```bash +rclone purge rustfs:rclone-demo +``` + +## Troubleshooting + +### `Access Denied` or empty listings + +Confirm `endpoint` includes the scheme and that the key pair matches a RustFS access key. The `region` value is required by the S3 signer even though RustFS ignores it; keep `us-east-1`. + +### Mount fails with `fusermount3: mount failed: Permission denied` + +Mounting needs the FUSE device and elevated privileges. Inside a container, run with `--device /dev/fuse --cap-add SYS_ADMIN` — and use `--privileged` if the mount helper still fails. On a host, verify that `fuse3` is installed and `/dev/fuse` exists. + +### Sync deleted nothing on the destination + +`rclone copy` never deletes. Only `rclone sync` (or `rclone delete`) removes destination objects, and `--dry-run` is the safe way to preview either. + +## Next steps + +- Review [S3 compatibility notes](/administration/protocols/s3) before adopting additional rclone backends. +- Create dedicated production credentials with [Access Key Management](/security-compliance/iam/access-token). +- Follow the [rclone S3 documentation](https://rclone.org/s3/) for flags such as `--transfers`, bandwidth limits, and crypt overlays. diff --git a/content/ja/developer/integration/others/tusd.md b/content/ja/developer/integration/others/tusd.md new file mode 100644 index 00000000..99f3c588 --- /dev/null +++ b/content/ja/developer/integration/others/tusd.md @@ -0,0 +1,170 @@ +--- +title: "tusd" +description: "Receive resumable uploads into RustFS with the tusd server's S3 backend." +--- + +This guide connects [tusd](https://github.com/tus/tusd) — the official reference implementation of the tus resumable-upload protocol — to **RustFS** as its S3 storage backend. You will run tusd against a RustFS bucket, create an upload with the tus protocol, send the file in two chunks with an interruption in between, resume from the reported offset, and verify the assembled object in the bucket. The workflow was verified with `tusproject/tusd:v2.10.1` against `rustfs/rustfs-x86-musl:v2.3.1`. + +You need Docker, or a local tusd binary. This deployment is intended for local integration testing, not production. + +## Architecture + +```mermaid +flowchart LR + Client["tus client"] -->|"POST / PATCH / HEAD"| tusd["tusd :8080"] + tusd -->|"multipart upload"| RustFS["RustFS :9000"] +``` + +tusd stores each in-progress upload as S3 multipart parts in the bucket. A client that loses its connection asks the server for the last committed offset with `HEAD` and continues from there — the data already received is never sent twice. + +## 1. Run tusd + +Create the bucket and start tusd with the S3 backend, replacing all connection placeholders. The AWS region must be set even though RustFS ignores it: + +```bash +rc mb rustfs/tus-uploads + +docker run -d --name tusd --network oo-rustfs_default -p 8080:8080 \ + -e AWS_ACCESS_KEY_ID= \ + -e AWS_SECRET_ACCESS_KEY= \ + -e AWS_REGION=us-east-1 \ + tusproject/tusd:latest \ + -s3-bucket tus-uploads \ + -s3-endpoint http://:9000 +``` + +Check that the server is healthy: + +```bash +curl -s -o /dev/null -w "%{http_code}\n" http://localhost:8080/health +``` + +```text +200 +``` + +## 2. Create the upload + +Create a 6 MiB upload and read the `Location` header: + +```bash +curl -s -D - -o /dev/null -X POST http://localhost:8080/files/ \ + -H "Upload-Length: 6291456" -H "Tus-Resumable: 1.0.0" \ + | grep -i "^Location:" +``` + +```text +Location: http://localhost:8080/files/b7338250daa9...+NGVmNDRhZjEt... +``` + +The upload URL contains the file ID and a message-authentication tag. Strip the scheme and host before re-sending it (the server echoes whatever `Host` it received, which may not be reachable from your next client). + +## 3. Upload in chunks with an interruption + +Send the first 2.5 MB, then stop — this is the point where a mobile client would lose its connection: + +```bash +head -c 2500000 demo.bin > part1.bin + +curl -s -o /dev/null -w "%{http_code}\n" -X PATCH "http://localhost:8080${LOC}" \ + -H "Upload-Offset: 0" -H "Tus-Resumable: 1.0.0" \ + -H "Content-Type: application/offset+octet-stream" \ + --data-binary @part1.bin +``` + +```text +204 +``` + +Ask the server how much it actually has — this is the resumable-upload core: + +```bash +curl -s -X HEAD "http://localhost:8080${LOC}" \ + -H "Tus-Resumable: 1.0.0" -D - -o /dev/null | grep -i upload-offset +``` + +```text +Upload-Offset: 2500000 +``` + +## 4. Resume and finish + +Continue from offset 2500000 with the remaining bytes: + +```bash +tail -c 3791456 demo.bin > part2.bin + +curl -s -o /dev/null -w "%{http_code}\n" -X PATCH "http://localhost:8080${LOC}" \ + -H "Upload-Offset: 2500000" -H "Tus-Resumable: 1.0.0" \ + -H "Content-Type: application/offset+octet-stream" \ + --data-binary @part2.bin +``` + +```text +204 +``` + +Download the finished upload through tusd and compare checksums with the source: + +```bash +curl -s -o download.bin "http://localhost:8080${LOC}" +sha1sum demo.bin download.bin +``` + +```text +d9016032ced6c7515b67a0c556e006c4b25a5858 demo.bin +d9016032ced6c7515b67a0c556e006c4b25a5858 download.bin +``` + +## 5. Verify objects in RustFS + +List the bucket: + +```bash +rc ls rustfs/tus-uploads/ -r +``` + +The bucket holds the assembled object plus one `.info` metadata file per upload — both live and finished: + +```text +b7338250daa9a1a79c1343502b57b28f 6 MiB +b7338250daa9a1a79c1343502b57b28f.info 378 B +``` + +The object key is the upload ID, and the object body is the uploaded file byte-for-byte — so any S3 client can read completed uploads directly from the bucket. + +![tus uploads stored in the RustFS Console](./images/rustfs-tus-uploads.png) + +## 6. Stop or reset + +To tear down the server while keeping the bucket objects: + +```bash +docker rm -f tusd +``` + +To delete the stored uploads: + +```bash +rc rm rustfs/tus-uploads/ --recursive --force +``` + +## Troubleshooting + +### `CreateMultipartUpload ... A region must be set when sending requests to S3` + +tusd builds its S3 client from the AWS environment, and the region is mandatory for endpoint resolution. Export `AWS_REGION=us-east-1` next to the credentials, as in step 1. + +### `PATCH` returns `404` or connects to the wrong host + +The `Location` URL echoes the `Host` header of the creation request. When your client and the server use different hostnames (container name versus published port), strip the scheme and host from the URL and send the path to the address the client can reach. + +### Upload disappears after server restart + +The S3 backend keeps `.info` files in the bucket, so uploads survive restarts. If you run tusd against an empty bucket that another process prunes, the metadata is lost — protect the `tus-uploads` prefix from cleanup jobs. + +## Next steps + +- Review [S3 compatibility notes](/administration/protocols/s3) before adopting additional tusd backends. +- Create dedicated production credentials with [Access Key Management](/security-compliance/iam/access-token). +- Follow the [tus protocol documentation](https://tus.io/protocols/resumable-upload) for creation-with-upload, termination, and checksum extensions that tusd supports on top of the core protocol. diff --git a/content/ja/developer/integration/reverse-proxy/envoy.md b/content/ja/developer/integration/reverse-proxy/envoy.md new file mode 100644 index 00000000..dc635a48 --- /dev/null +++ b/content/ja/developer/integration/reverse-proxy/envoy.md @@ -0,0 +1,192 @@ +--- +title: "Envoy" +description: "Deploy RustFS behind Envoy with TLS-terminated routes for the S3 API and Console." +--- + +Use **Envoy** to terminate TLS and route separate hostnames to the RustFS S3 API and Console. This deployment runs Envoy and a single-node RustFS instance on one Docker network. You need Docker Engine, two DNS records, and a TLS certificate that covers both hostnames (the guide uses a self-signed certificate for testing). + +This guide uses these example hostnames: + +- `s3.example.com` for the S3 API +- `console.example.com` for the Console + +Replace them with hostnames that resolve to the Docker host. + +:::warning[Serve S3 from the root path] + +Do not publish the S3 API under a path such as `/s3/`. AWS Signature Version 4 includes the request path and host, so rewriting either value can invalidate signed requests. Envoy forwards the incoming `Host` header unchanged, which keeps signatures valid. + +::: + +## 1. Create the deployment directories + +Create directories for the Envoy configuration and TLS certificate: + +```bash +mkdir -p rustfs-envoy/certs +cd rustfs-envoy +``` + +For local testing, generate a self-signed certificate covering both hostnames: + +```bash +openssl req -x509 -newkey rsa:2048 -nodes \ + -keyout certs/privkey.pem -out certs/fullchain.pem -days 30 \ + -subj "/CN=*.example.com" \ + -addext "subjectAltName=DNS:s3.example.com,DNS:console.example.com" +chmod 644 certs/privkey.pem +``` + +Make the certificate readable by the non-root user the Envoy image runs as — a `600` private key produces a misleading `Failed to load incomplete private key` error. + +## 2. Configure Envoy + +Create the configuration with an HTTPS listener and two virtual hosts. Route timeouts are disabled (`timeout: 0s`) so long streaming S3 uploads are not cut off: + +```yaml title="envoy.yaml" +static_resources: + listeners: + - name: https + address: {socket_address: {address: 0.0.0.0, port_value: 8443}} + filter_chains: + - transport_socket: + name: envoy.transport_sockets.tls + typed_config: + "@type": type.googleapis.com/envoy.extensions.transport_sockets.tls.v3.DownstreamTlsContext + common_tls_context: + tls_certificates: + - certificate_chain: {filename: /certs/fullchain.pem} + private_key: {filename: /certs/privkey.pem} + filters: + - name: envoy.filters.network.http_connection_manager + typed_config: + "@type": type.googleapis.com/envoy.extensions.filters.network.http_connection_manager.v3.HttpConnectionManager + stat_prefix: rustfs_https + route_config: + virtual_hosts: + - name: s3 + domains: ["s3.example.com", "s3.example.com:*"] + routes: + - match: {prefix: "/"} + route: {cluster: rustfs_s3, timeout: 0s} + - name: console + domains: ["console.example.com", "console.example.com:*"] + routes: + - match: {prefix: "/"} + route: {cluster: rustfs_console, timeout: 0s} + http_filters: + - name: envoy.filters.http.router + typed_config: + "@type": type.googleapis.com/envoy.extensions.filters.http.router.v3.Router + clusters: + - name: rustfs_s3 + connect_timeout: 5s + type: STRICT_DNS + lb_policy: ROUND_ROBIN + load_assignment: + cluster_name: rustfs_s3 + endpoints: + - lb_endpoints: + - endpoint: {address: {socket_address: {address: rustfs, port_value: 9000}}} + - name: rustfs_console + connect_timeout: 5s + type: STRICT_DNS + lb_policy: ROUND_ROBIN + load_assignment: + cluster_name: rustfs_console + endpoints: + - lb_endpoints: + - endpoint: {address: {socket_address: {address: rustfs, port_value: 9001}}} +``` + +The `host:*` domain entries matter: clients send `Host: s3.example.com:8443` on non-standard ports, and Envoy matches the authority including the port. + +## 3. Start Envoy + +Run Envoy on the same Docker network as RustFS, publishing only the proxy port: + +```bash +docker run -d --name envoy --network oo-rustfs_default -p 8443:8443 \ + -v "$PWD/envoy.yaml":/envoy.yaml:ro \ + -v "$PWD/certs":/certs:ro \ + envoyproxy/envoy:v1.34-latest -c /envoy.yaml +``` + +## 4. Verify both endpoints + +Point the example hostnames at the proxy with `curl --resolve` (in production, DNS does this): + +```bash +curl -sk --resolve s3.example.com:8443:127.0.0.1 \ + https://s3.example.com:8443/health/ready -o /dev/null -w "s3 api: %{http_code}\n" + +curl -sk --resolve console.example.com:8443:127.0.0.1 \ + https://console.example.com:8443/rustfs/console/ -o /dev/null -w "console: %{http_code}\n" +``` + +```text +s3 api: 200 +console: 200 +``` + +`-k` skips certificate validation because the certificate is self-signed; with a trusted certificate, drop it. + +## 5. Send signed S3 requests through Envoy + +Point any S3 client at the proxy as if it were RustFS. Configure the client with `https://s3.example.com:8443` as the endpoint and path-style addressing; when the proxy certificate is trusted, signed AWS Signature Version 4 requests pass through unchanged. For a quick test over plain HTTP, add an HTTP listener on port 8080 with the same virtual-host routing as the HTTPS listener, then use the endpoint `http://s3.example.com:8080`: + +```bash +rc alias set rustfs-envoy http://s3.example.com:8080 +rc ls rustfs-envoy/rclone-demo/ +``` + +```text +[ ] 0B seed/ +[2026-09-29 11:12:41] 5 B write-test.txt +``` + +The signatures validate because Envoy forwards the original `Host` header to RustFS. + +## Multi-node backends + +For a distributed RustFS deployment, add every node to the S3 cluster: + +```yaml title="envoy.yaml" + - name: rustfs_s3 + connect_timeout: 5s + type: STRICT_DNS + lb_policy: ROUND_ROBIN + load_assignment: + cluster_name: rustfs_s3 + endpoints: + - lb_endpoints: + - endpoint: {address: {socket_address: {address: node1, port_value: 9000}}} + - lb_endpoints: + - endpoint: {address: {socket_address: {address: node2, port_value: 9000}}} + - lb_endpoints: + - endpoint: {address: {socket_address: {address: node3, port_value: 9000}}} + - lb_endpoints: + - endpoint: {address: {socket_address: {address: node4, port_value: 9000}}} +``` + +The Console cluster follows the same pattern on port `9001`. + +## Troubleshooting + +### `Failed to load incomplete private key from path` + +The Envoy container runs as a non-root user and cannot read a `600` root-owned key. `chmod 644` the key files (or chown them to the container user, UID `1001` in the official image). + +### Routes return `404` with the correct hostnames + +Envoy matches the authority including the port. Add the `host:*` variants to each virtual host's `domains` list, as in the configuration above. + +### `Access Denied` from RustFS on proxied requests + +Confirm the proxy is not rewriting the path or the `Host` header. Signed requests must reach RustFS with the host the client signed for. + +## Next steps + +- [Configure an S3 client](/developer/examples/aws-cli) +- [Enable virtual-hosted-style bucket URLs](/integration/virtual) +- [Review health and readiness endpoints](/operations/status-check) diff --git a/content/ja/developer/integration/reverse-proxy/index.md b/content/ja/developer/integration/reverse-proxy/index.md index eff7324d..397cdad6 100644 --- a/content/ja/developer/integration/reverse-proxy/index.md +++ b/content/ja/developer/integration/reverse-proxy/index.md @@ -13,6 +13,7 @@ We recommend using separate hostnames for the S3 API on port `9000` and the Cons - [Traefik](./traefik.md) - [Caddy](./caddy.md) - [HAProxy](./haproxy.md) +- [Envoy](./envoy.md) - [Apache HTTP Server](./httpd.md) ## Related configuration diff --git a/content/ja/developer/integration/reverse-proxy/meta.json b/content/ja/developer/integration/reverse-proxy/meta.json index f3cbddfe..93f3e6df 100644 --- a/content/ja/developer/integration/reverse-proxy/meta.json +++ b/content/ja/developer/integration/reverse-proxy/meta.json @@ -4,7 +4,8 @@ "nginx", "traefik", "caddy", + "envoy", "haproxy", "httpd" ] -} \ No newline at end of file +} diff --git a/content/zh/developer/integration/ai/images/rustfs-vllm-models.png b/content/zh/developer/integration/ai/images/rustfs-vllm-models.png new file mode 100644 index 00000000..fa25d229 Binary files /dev/null and b/content/zh/developer/integration/ai/images/rustfs-vllm-models.png differ diff --git a/content/zh/developer/integration/ai/index.md b/content/zh/developer/integration/ai/index.md index 00a55f9d..b46c73e5 100644 --- a/content/zh/developer/integration/ai/index.md +++ b/content/zh/developer/integration/ai/index.md @@ -8,5 +8,6 @@ description: "通过 S3 兼容的对象存储接口,将 AI 平台连接到 Rus ## 平台 - [Ray](./ray.md) +- [vLLM](./vllm.md) 请使用专用的存储桶保存训练数据集与检查点,并为凭证仅授予所需桶操作的权限。 diff --git a/content/zh/developer/integration/ai/meta.json b/content/zh/developer/integration/ai/meta.json index 573fc520..b30f9e0f 100644 --- a/content/zh/developer/integration/ai/meta.json +++ b/content/zh/developer/integration/ai/meta.json @@ -1,6 +1,7 @@ { "title": "AI", "pages": [ - "ray" + "ray", + "vllm" ] } diff --git a/content/zh/developer/integration/ai/vllm.md b/content/zh/developer/integration/ai/vllm.md new file mode 100644 index 00000000..b8bcb13a --- /dev/null +++ b/content/zh/developer/integration/ai/vllm.md @@ -0,0 +1,154 @@ +--- +title: "vLLM" +description: "用 vLLM 加载存放在 RustFS 中的模型权重进行 LLM 推理服务。" +--- + +本指南将高吞吐 LLM 推理引擎 [vLLM](https://github.com/vllm-project/vllm) 连接到 **RustFS** 作为其模型权重存储。你将把模型上传到 RustFS 桶,通过 rclone 挂载把桶暴露给 vLLM 宿主机,并以 OpenAI 兼容 API 提供服务。整个流程使用 `vllm/vllm-openai-cpu`(vLLM 0.30.0)在纯 CPU 主机上对存放在 `rustfs/rustfs-x86-musl:v2.3.1` 桶中的 `facebook/opt-125m` 验证通过。 + +你需要 Docker 以及运行 vLLM 的宿主机上的 rclone 二进制。本部署用于本地集成测试,不适用于生产环境。 + +## 架构 + +```mermaid +flowchart LR + Upload["rclone copy"] -->|"weights"| RustFS["RustFS :9000"] + RustFS -->|"rclone mount"| Mount["/mnt/vllm-models"] + Mount -->|"weight load"| vLLM["vLLM :8000"] + Client["OpenAI SDK / curl"] -->|"completions"| vLLM +``` + +桶内只有这一份模型。各服务节点以只读方式挂载桶,权重统一从 RustFS 拉取,本地不存在会漂移的模型副本。 + +:::note[为什么用 rclone 挂载] + +vLLM 0.30 通过 RunAI model streamer 加载 `s3://` 模型路径,其对自定义 S3 端点的分段读取目前无法工作(任何非零偏移都会报 `File access error`)。把桶挂载为文件系统是经过验证的方式——权重保留在 RustFS,vLLM 按本地文件读取。 + +::: + +## 1. 上传模型到 RustFS + +创建桶并把模型权重拷入,替换全部连接占位符: + +```ini title="rclone.conf" +[rustfs] +type = s3 +provider = Other +access_key_id = +secret_access_key = +endpoint = http://:9000 +region = us-east-1 +``` + +```bash +rc mb rustfs/vllm-models +rclone copy ./opt-125m rustfs:vllm-models/opt-125m --transfers 4 +``` + +任何 Hugging Face 目录结构都可以——`config.json`、tokenizer 文件与权重文件(`model.safetensors` 或 `pytorch_model.bin`)。每个模型使用一个独立前缀,多个模型即可共用同一个桶。 + +## 2. 在 vLLM 宿主机上挂载桶 + +在运行 vLLM 的机器上,以 `--allow-other` 只读挂载桶供其他用户访问: + +```bash +mkdir -p /mnt/vllm-models +rclone mount rustfs:vllm-models /mnt/vllm-models \ + --allow-other --daemon +ls /mnt/vllm-models/opt-125m/ +``` + +```text +config.json merges.txt model.safetensors tokenizer.json vocab.json +``` + +## 3. 运行 vLLM + +用 CPU 镜像指向挂载的权重启动: + +```bash +docker run -d --name vllm -p 8000:8000 --shm-size=2g \ + -v /mnt/vllm-models:/models:ro \ + vllm/vllm-openai-cpu:latest \ + --model /models/opt-125m --served-model-name opt-125m \ + --dtype float32 --max-model-len 256 --gpu-memory-utilization 0.15 +``` + +vLLM 经由挂载读取权重——容器本身无状态,模型保存在 RustFS。等待服务就绪: + +```bash +curl -s http://localhost:8000/v1/models | head -c 200 +``` + +```text +{"object":"list","data":[{"id":"opt-125m","object":"model","created":...,"root":"/models/opt-125m",...}]} +``` + +在 CPU 后端上 `--gpu-memory-utilization` 控制预留给 KV cache 的内存比例;小内存主机请调低。`--dtype float32` 与该模型在 CPU attention 内核上支持的精度匹配。 + +## 4. 推理 + +发送 OpenAI 兼容的 completion 请求: + +```bash +curl -s http://localhost:8000/v1/completions \ + -H "Content-Type: application/json" \ + -d '{"model": "opt-125m", "prompt": "RustFS is", "max_tokens": 12, "temperature": 0}' +``` + +```json +{"id":"cmpl-...","object":"text_completion","model":"opt-125m", + "choices":[{"index":0,"text":" a great tool for building your own server. It's a", + "finish_reason":"length",...}]} +``` + +请求是标准 OpenAI 结构,Python 客户端无需改动: + +```python +from openai import OpenAI + +client = OpenAI(base_url="http://localhost:8000/v1", api_key="EMPTY") +print(client.completions.create( + model="opt-125m", prompt="RustFS is", max_tokens=12, temperature=0, +).choices[0].text) +``` + +![存储在 RustFS 控制台中的 vLLM 模型权重](./images/rustfs-vllm-models.png) + +## 5. 停止或重置 + +保留桶内对象、仅拆除演示环境: + +```bash +docker rm -f vllm +fusermount -u /mnt/vllm-models +``` + +删除已存储的模型: + +```bash +rclone purge rustfs:vllm-models +``` + +## 故障排查 + +### `Cannot find any model weights with /models/...` + +挂载的目录缓存过期,或权重根本没传到桶里。重新执行 `rclone copy`,并在启动 vLLM 前先用 `ls` 通过挂载点确认文件存在。迭代调试时给挂载加较短的 `--dir-cache-time`(如 `10s`)。 + +### `Unsupported CPU attention configuration: head_dim=...` + +vLLM 的 CPU 内核只支持有限的 head_dim 集合。`hf-internal-testing/tiny-random-*` 这类微型测试模型的奇异形状会在请求阶段失败——请使用真实的小模型,如 `facebook/opt-125m`。 + +### `Insufficient space in /dev/shm` + +vLLM 的 CPU 引擎通过共享内存交换张量。运行容器时加 `--shm-size=2g`(或 `--ipc=host`)。 + +### 服务器退出并提示 `Available memory on node 0 ... is less than desired CPU memory utilization` + +默认的 KV cache 预留是系统内存的 90%。用 `--gpu-memory-utilization 0.15` 调低(该 flag 在 CPU 后端同样生效,虽然名字里带 GPU)。 + +## 下一步 + +- 在启用更多服务方案前,先阅读 [S3 兼容性说明](/administration/protocols/s3)。 +- 使用[访问密钥管理](/security-compliance/iam/access-token)创建专用的生产凭证。 +- 按照 [vLLM 文档](https://docs.vllm.ai/en/latest/serving/openai_compatible_server.html)在同一个桶支撑的模型库之上使用 chat 模板、张量并行与量化权重。 diff --git a/content/zh/developer/integration/big-data/airflow.md b/content/zh/developer/integration/big-data/airflow.md new file mode 100644 index 00000000..43c8dbaa --- /dev/null +++ b/content/zh/developer/integration/big-data/airflow.md @@ -0,0 +1,170 @@ +--- +title: "Airflow" +description: "通过 Amazon S3 provider 在 Airflow DAG 与 RustFS 之间移动数据。" +--- + +本指南将工作流编排平台 [Apache Airflow](https://github.com/apache/airflow) 通过 Amazon S3 provider 的 hook、operator 与 sensor 连接到 **RustFS**。你将注册一个自定义端点的连接,运行一个向 RustFS 桶写入对象的 DAG,用 `S3KeySensor` 等待键出现,再用 `S3Hook` 读回对象。整个流程使用 `apache/airflow:3.3.2`(standalone,SequentialExecutor)和 `apache-airflow-providers-amazon` provider 对 `rustfs/rustfs-x86-musl:v2.3.1` 验证通过。 + +你需要安装 Docker,或一个现有的 Airflow 实例。本部署用于本地集成测试,不适用于生产环境。 + +## 架构 + +```mermaid +flowchart LR + Scheduler["Airflow scheduler"] -->|"tasks"| Hook["S3Hook / operators"] + Hook -->|"S3 API"| RustFS["RustFS :9000"] + Sensor["S3KeySensor"] -->|"poll key"| RustFS +``` + +DAG 内所有 S3 交互都经过 provider 的 S3 客户端,由连接的 `endpoint_url` 指向 RustFS。operator、sensor 与 hook 共享同一个连接对象。 + +## 1. 运行 Airflow + +以关闭示例的方式启动 standalone 实例,并创建演示桶: + +```bash +docker run -d --name airflow --network oo-rustfs_default -p 8080:8080 \ + -e AIRFLOW__CORE__LOAD_EXAMPLES=False \ + -v "$PWD/dags":/opt/airflow/dags \ + apache/airflow:3.3.2 standalone + +rc mb rustfs/airflow-demo +``` + +镜像预装了全部 provider,包括 `apache-airflow-providers-amazon`。 + +## 2. 注册 RustFS 连接 + +S3 provider 从连接的 extra 字段读取端点。替换全部连接占位符: + +```bash +docker exec airflow airflow connections add rustfs \ + --conn-type aws \ + --conn-extra '{"endpoint_url": "http://:9000", "region_name": "us-east-1", "aws_access_key_id": "", "aws_secret_access_key": ""}' +``` + +`--conn-extra` 中的 `aws_access_key_id` 与 `aws_secret_access_key` 提供凭证;`endpoint_url` 把 boto3 客户端从 AWS 重定向到 RustFS。 + +## 3. 编写 DAG + +该 DAG 用 operator 写入对象、用 sensor 等待键、再用 hook 读回: + +```python title="rustfs_demo.py" +import datetime + +from airflow.providers.amazon.aws.hooks.s3 import S3Hook +from airflow.providers.amazon.aws.operators.s3 import S3CreateObjectOperator +from airflow.providers.amazon.aws.sensors.s3 import S3KeySensor +from airflow.sdk import dag, task + +@dag( + schedule=None, + start_date=datetime.datetime(2026, 1, 1), + catchup=False, + tags=["rustfs"], +) +def rustfs_demo(): + create = S3CreateObjectOperator( + task_id="write_object", + s3_bucket="airflow-demo", + s3_key="dags/airflow-put.txt", + data="written by airflow to rustfs", + aws_conn_id="rustfs", + replace=True, + ) + + wait = S3KeySensor( + task_id="wait_for_object", + bucket_key="dags/airflow-put.txt", + bucket_name="airflow-demo", + aws_conn_id="rustfs", + timeout=120, + poke_interval=10, + mode="reschedule", + ) + + @task + def read_object(): + hook = S3Hook(aws_conn_id="rustfs") + body = hook.read_key(key="dags/airflow-put.txt", bucket_name="airflow-demo") + print("read back:", body) + assert body == "written by airflow to rustfs" + + create >> [wait, read_object()] + +rustfs_demo() +``` + +注意导入路径:`S3CreateObjectOperator` 位于 `operators` 模块,而 `S3KeySensor` 位于 `sensors` 模块——从同一个模块导入两者会失败。 + +## 4. 取消暂停并触发 + +新 DAG 默认处于暂停状态,暂停期间触发的 run 会永远停留在排队状态。先取消暂停,再触发: + +```bash +docker exec airflow airflow dags unpause rustfs_demo +docker exec airflow airflow dags trigger rustfs_demo +``` + +观察 run 完成: + +```bash +docker exec airflow airflow dags list-runs rustfs_demo | head -3 +``` + +```text +dag_id run_id state +rustfs_demo manual__2026-09-29T13:15:57.332688+00:00 success +``` + +三个任务全部成功:`write_object`、`wait_for_object` 和 `read_object`。 + +## 5. 验证 RustFS 中的对象 + +列举桶内前缀: + +```bash +rc ls rustfs/airflow-demo/ -r +rc cat rustfs/airflow-demo/dags/airflow-put.txt +``` + +```text +[2026-09-29 13:16:01] 28 B dags/airflow-put.txt +written by airflow to rustfs +``` + +![存储在 RustFS 控制台中的 Airflow 对象](./images/rustfs-airflow-object.png) + +## 6. 停止或重置 + +保留桶内对象、仅拆除演示环境: + +```bash +docker rm -f airflow +``` + +删除已存储的数据: + +```bash +rc rm rustfs/airflow-demo/ --recursive --force +``` + +## 故障排查 + +### 触发后 dag run 一直停留在 `queued` + +DAG 处于暂停状态。Airflow 3 中新 DAG 默认暂停,此状态下触发的 run 永远不会执行。执行 `airflow dags unpause rustfs_demo`,排队的 run 会自行启动。 + +### `cannot import name 'S3KeySensor' from 'airflow.providers.amazon.aws.operators.s3'` + +sensor 在独立的模块里:`from airflow.providers.amazon.aws.sensors.s3 import S3KeySensor`。 + +### 任务报连接错误 + +`endpoint_url` 必须从 Airflow 容器内部可达——请使用 RustFS 的 Docker 网络主机名,而不是 `localhost`。需要检查组件状态时,Airflow 3 的健康检查端点在 `/api/v2/monitor/health`。 + +## 下一步 + +- 在启用更多 provider hook 前,先阅读 [S3 兼容性说明](/administration/protocols/s3)。 +- 使用[访问密钥管理](/security-compliance/iam/access-token)创建专用的生产凭证。 +- 按照 [Amazon provider 文档](https://airflow.apache.org/docs/apache-airflow-providers-amazon/stable/index.html)了解 `S3ToLocalFilesystemOperator`、`LocalFilesystemToS3Operator` 等传输 operator。 diff --git a/content/zh/developer/integration/big-data/delta-lake.md b/content/zh/developer/integration/big-data/delta-lake.md new file mode 100644 index 00000000..fd966b50 --- /dev/null +++ b/content/zh/developer/integration/big-data/delta-lake.md @@ -0,0 +1,125 @@ +--- +title: "Delta Lake" +description: "用 delta-rs 在 RustFS 上读写 Delta 表。" +--- + +本指南将开源湖仓表格式 [Delta Lake](https://github.com/delta-io/delta) 通过 Rust 原生的 Delta 实现 delta-rs 连接到 **RustFS**。你将从 Python 向 RustFS 桶写入一张 Delta 表,带着 ACID 事务历史读回,并确认桶内的 `_delta_log` 与 Parquet 文件。整个流程使用 `deltalake` Python 包(delta-rs)和 pandas 对 `rustfs/rustfs-x86-musl:v2.3.1` 验证通过。 + +你需要 Python 3.9 或更高版本。本部署用于本地集成测试,不适用于生产环境。 + +## 架构 + +```mermaid +flowchart LR + DF["pandas DataFrame"] -->|"write_deltalake"| deltaRS["delta-rs"] + deltaRS -->|"Parquet + _delta_log"| RustFS["RustFS :9000"] + Query["DeltaTable"] -->|"read / time travel"| RustFS +``` + +delta-rs 把每张表存为 Parquet 文件加事务日志(`_delta_log/`)。所有 I/O 都经过 `object_store` crate,与其他 S3 客户端一样使用 AWS 环境变量配置。 + +## 1. 安装客户端 + +```bash +pip install deltalake pandas pyarrow +``` + +`pyarrow` 负责把 pandas DataFrame 转换成 Delta 兼容的 record batch。 + +## 2. 写入 Delta 表 + +创建桶并写入一张表,替换全部连接占位符。由于 RustFS 不提供 copy-if-not-exists(delta-rs 处理提交冲突时依赖它),需要设置 `AWS_S3_ALLOW_UNSAFE_RENAME`: + +```python title="delta_s3.py" +import pandas as pd +from deltalake import DeltaTable, write_deltalake + +storage_options = { + "AWS_ENDPOINT_URL": "http://:9000", + "AWS_ACCESS_KEY_ID": "", + "AWS_SECRET_ACCESS_KEY": "", + "AWS_REGION": "us-east-1", + "AWS_ALLOW_HTTP": "true", + "AWS_S3_ALLOW_UNSAFE_RENAME": "true", +} + +table = "s3:///events" +df = pd.DataFrame({"id": [1, 2, 3], "name": ["alpha", "beta", "gamma"]}) +write_deltalake(table, df, storage_options=storage_options) +print("written:", df.shape[0], "rows") +``` + +```text +written: 3 rows +``` + +`table` URI 使用标准的 `s3://bucket/prefix` 形式;端点与凭证来自 `storage_options`。 + +## 3. 读回表 + +```python title="delta_read.py" +from deltalake import DeltaTable + +back = DeltaTable("s3:///events", storage_options=storage_options).to_pandas() +print("read back:", back.shape[0], "rows") +print(back.sort_values("id").to_string(index=False)) +print("version:", DeltaTable("s3:///events", storage_options=storage_options).version()) +``` + +```text +read back: 3 rows + id name + 1 alpha + 2 beta + 3 gamma +version: 0 +``` + +版本记录在事务日志中,因此同一张表支持 `DeltaTable(..., version=N)` 做时间旅行,追加写入会递增版本号。 + +## 4. 验证 RustFS 中的对象 + +列举表前缀: + +```bash +rc ls rustfs// -r +``` + +首次提交会创建事务日志和一个 Parquet 文件: + +```text +events/_delta_log/00000000000000000000.json +events/part-00000-3859855e-45e4-4ae5-94ff-2d8eab5e7ebb-c000.snappy.parquet +``` + +每次新写入都会新增一个 `NNNNNNNNNNNNNNNNNNNN.json` 日志条目和 Parquet 分件;读取方重放日志即可得到一致的快照。 + +![存储在 RustFS 控制台中的 Delta 表文件](./images/rustfs-delta-table.png) + +## 5. 停止或重置 + +delta-rs 自身不持有状态。删除表: + +```bash +rc rm rustfs//events/ --recursive --force +``` + +## 故障排查 + +### 写入 pandas DataFrame 时报 `Import pyarrow failed` + +`write_deltalake` 经由 Arrow 转换 DataFrame。请与 `deltalake`、`pandas` 一并安装 `pyarrow`。 + +### 提交时报 `Generic DeltaTable error: commit conflict` 或 rename 错误 + +delta-rs 通过复制并重命名临时对象来提交,这要求后端支持原子重命名。对不支持 copy-if-not-exists 的 S3 兼容存储,需在 `storage_options` 中设置 `AWS_S3_ALLOW_UNSAFE_RENAME: "true"`——单写者可接受,并发写者不适用。 + +### `Unknown lengthy error: AWS connectivity or endpoint errors` + +确认 `AWS_ENDPOINT_URL` 带协议前缀,且纯 HTTP 端点把 `AWS_ALLOW_HTTP` 设置为 `"true"`;否则 S3 客户端只讲 HTTPS。 + +## 下一步 + +- 在启用更多 Delta 客户端前,先阅读 [S3 兼容性说明](/administration/protocols/s3)。 +- 使用[访问密钥管理](/security-compliance/iam/access-token)创建专用的生产凭证。 +- 按照 [delta-rs 使用文档](https://delta-io.github.io/delta-rs/usage/writing/writing-to-s3/)了解基于 DynamoDB 提交协调的并发写者方案。 diff --git a/content/zh/developer/integration/big-data/images/rustfs-airflow-object.png b/content/zh/developer/integration/big-data/images/rustfs-airflow-object.png new file mode 100644 index 00000000..03dfbd63 Binary files /dev/null and b/content/zh/developer/integration/big-data/images/rustfs-airflow-object.png differ diff --git a/content/zh/developer/integration/big-data/images/rustfs-delta-table.png b/content/zh/developer/integration/big-data/images/rustfs-delta-table.png new file mode 100644 index 00000000..5c753260 Binary files /dev/null and b/content/zh/developer/integration/big-data/images/rustfs-delta-table.png differ diff --git a/content/zh/developer/integration/big-data/images/rustfs-kafka-sink.png b/content/zh/developer/integration/big-data/images/rustfs-kafka-sink.png new file mode 100644 index 00000000..bae09a03 Binary files /dev/null and b/content/zh/developer/integration/big-data/images/rustfs-kafka-sink.png differ diff --git a/content/zh/developer/integration/big-data/index.md b/content/zh/developer/integration/big-data/index.md index a8a56f91..eb84c3b2 100644 --- a/content/zh/developer/integration/big-data/index.md +++ b/content/zh/developer/integration/big-data/index.md @@ -8,6 +8,7 @@ description: "通过 S3 兼容的对象存储接口将数据分析系统连接 ## 系统 - [ClickHouse](./clickhouse.md) +- [Airflow](./airflow.md) - [Hudi](./hudi.md) - [Iceberg](./iceberg.md) - [PyIceberg](./pyiceberg.md) @@ -16,8 +17,10 @@ description: "通过 S3 兼容的对象存储接口将数据分析系统连接 - [OpenDAL](./opendal.md) - [DuckDB](./duckdb.md) - [Doris](./doris.md) +- [Delta Lake](./delta-lake.md) - [lakeFS](./lakefs.md) - [InfluxDB](./influxdb.md) +- [Kafka](./kafka.md) - [Spark](./spark.md) - [Flink](./flink.md) - [Trino](./trino.md) diff --git a/content/zh/developer/integration/big-data/kafka.md b/content/zh/developer/integration/big-data/kafka.md new file mode 100644 index 00000000..7cd4b0ed --- /dev/null +++ b/content/zh/developer/integration/big-data/kafka.md @@ -0,0 +1,178 @@ +--- +title: "Kafka" +description: "通过 Kafka Connect S3 sink 连接器把 Kafka 主题数据下沉到 RustFS。" +--- + +本指南将分布式事件流平台 [Apache Kafka](https://github.com/apache/kafka) 通过 Kafka Connect S3 sink 连接器连接到 **RustFS**。你将运行一个 KRaft broker 和一个 Connect worker,为一个主题部署 S3 sink,并生产一批落为桶内对象的记录。整个流程使用 `apache/kafka:4.0.0` 和 `confluentinc/kafka-connect-s3` v10.5.25 对 `rustfs/rustfs-x86-musl:v2.3.1` 验证通过。 + +Kafka 的 KIP-405 分层存储需要 `RemoteLogStorageManager` 插件,而生态中的 S3 实现均为厂商专有。Connect S3 sink 是把主题数据迁移到 S3 兼容存储的开放、可自托管方案,也是本指南采用的方式。 + +你需要安装 Docker。本部署用于本地集成测试,不适用于生产环境。 + +## 架构 + +```mermaid +flowchart LR + Producer["Console producer"] -->|"records"| Broker["Kafka broker :9092"] + Broker -->|"consumer group"| Connect["Connect S3 sink"] + Connect -->|"batched objects"| RustFS["RustFS :9000"] +``` + +sink 任务以独立的消费者组消费主题,并按每个分区每 `flush.size` 条记录一个对象的方式写入桶内。 + +## 1. 运行 broker + +启动一个发布地址对其他容器可达的 KRaft broker: + +```bash +docker run -d --name kafka --hostname kafka --network oo-rustfs_default \ + -e CLUSTER_ID=5L6g3nShT-eMCtK--X86sw \ + -e KAFKA_NODE_ID=1 \ + -e KAFKA_PROCESS_ROLES=broker,controller \ + -e KAFKA_LISTENERS=PLAINTEXT://:9092,CONTROLLER://:9093 \ + -e KAFKA_ADVERTISED_LISTENERS=PLAINTEXT://kafka:9092 \ + -e KAFKA_CONTROLLER_LISTENER_NAMES=CONTROLLER \ + -e KAFKA_LISTENER_SECURITY_PROTOCOL_MAP=CONTROLLER:PLAINTEXT,PLAINTEXT:PLAINTEXT \ + -e KAFKA_CONTROLLER_QUORUM_VOTERS=1@kafka:9093 \ + -e KAFKA_OFFSETS_TOPIC_REPLICATION_FACTOR=1 \ + -e KAFKA_TRANSACTION_STATE_LOG_REPLICATION_FACTOR=1 \ + -e KAFKA_TRANSACTION_STATE_LOG_MIN_ISR=1 \ + apache/kafka:4.0.0 + +docker exec kafka /opt/kafka/bin/kafka-topics.sh \ + --bootstrap-server localhost:9092 \ + --create --topic rustfs-topic --partitions 1 --replication-factor 1 + +rc mb rustfs/kafka-demo +``` + +镜像默认发布 `localhost:9092`,仅在 broker 容器内部可用。`KAFKA_ADVERTISED_LISTENERS` 覆盖是让 Connect(以及任何远程客户端)能够连上 broker 的关键。 + +## 2. 安装连接器 + +下载打包了连接器及其依赖的 Confluent Hub 压缩包,解压到 worker 可见的目录: + +```bash +curl -Lo kafka-connect-s3.zip "https://hub-downloads.confluent.io/api/plugins/confluentinc/kafka-connect-s3/versions/10.5.25/confluentinc-kafka-connect-s3-10.5.25.zip" +unzip kafka-connect-s3.zip -d /opt/kafka-conn/plugins +``` + +## 3. 配置 worker 与 sink + +创建 worker 配置。value converter 必须是 `ByteArrayConverter`,记录才会按原始字节写出: + +```ini title="worker.properties" +bootstrap.servers=kafka:9092 +key.converter=org.apache.kafka.connect.storage.StringConverter +value.converter=org.apache.kafka.connect.converters.ByteArrayConverter +offset.storage.file.filename=/tmp/connect.offsets +offset.flush.interval.ms=5000 +plugin.path=/opt/kafka-conn-plugins +``` + +创建 sink 连接器配置,并替换全部连接占位符: + +```ini title="rustfs-sink.properties" +name=rustfs-sink +connector.class=io.confluent.connect.s3.S3SinkConnector +tasks.max=1 +topics=rustfs-topic +s3.bucket.name=kafka-demo +s3.region=us-east-1 +store.url=http://:9000 +s3.path.style.access.enabled=true +flush.size=3 +storage.class=io.confluent.connect.s3.storage.S3Storage +format.class=io.confluent.connect.s3.format.bytearray.ByteArrayFormat +consumer.override.auto.offset.reset=earliest +``` + +Kafka 4.0 把该类移动到了 `org.apache.kafka.connect.converters.ByteArrayConverter`——旧的 `storage` 包路径已无法解析。 + +## 4. 运行 worker + +以前台方式运行 `connect-standalone`,让日志进入 `docker logs`,桶凭证走环境变量: + +```bash +docker run -d --name kafka-connect --hostname kafka-connect \ + --network oo-rustfs_default \ + -v /opt/kafka-conn/plugins:/opt/kafka-conn-plugins:ro \ + -v "$PWD/worker.properties":/etc/kafka/worker.properties:ro \ + -v "$PWD/rustfs-sink.properties":/etc/kafka/sink.properties:ro \ + -e AWS_ACCESS_KEY_ID= \ + -e AWS_SECRET_ACCESS_KEY= \ + apache/kafka:4.0.0 \ + /opt/kafka/bin/connect-standalone.sh /etc/kafka/worker.properties /etc/kafka/sink.properties +``` + +当 sink 任务认领分区后,worker 即就绪: + +```text +INFO [rustfs-sink|task-0] Assigned topic partitions: [rustfs-topic-0] +``` + +## 5. 生产记录 + +至少发送 `flush.size` 条记录,让连接器完成一个批次: + +```bash +docker exec kafka sh -c "printf 'msg-one\nmsg-two\nmsg-three\n' | \ + /opt/kafka/bin/kafka-console-producer.sh --bootstrap-server localhost:9092 --topic rustfs-topic" +``` + +几秒后,这一批记录变成桶里的对象: + +```bash +rc ls rustfs/kafka-demo/ -r +rc cat rustfs/kafka-demo/topics/rustfs-topic/partition=0/rustfs-topic+0+0000000000.bin +``` + +```text +topics/rustfs-topic/partition=0/rustfs-topic+0+0000000000.bin +topics/rustfs-topic/partition=0/rustfs-topic+0+0000000003.bin +msg-one +msg-two +msg-three +``` + +对象名编码了主题、分区和起始偏移量。后续每三行记录落入下一个对象(`+0000000003.bin`,以此类推)。 + +![存储在 RustFS 控制台中的 Kafka sink 对象](./images/rustfs-kafka-sink.png) + +## 6. 停止或重置 + +保留桶内对象、仅拆除演示环境: + +```bash +docker rm -f kafka-connect kafka +``` + +删除已存储的数据: + +```bash +rc rm rustfs/kafka-demo/ --recursive --force +``` + +## 故障排查 + +### `AdminClient ... Rebootstrapping with Cluster (id: null)` 无限循环 + +broker 发布的是 `localhost:9092`,远程客户端收到的元数据指向它自己。按第 1 步设置 `KAFKA_ADVERTISED_LISTENERS=PLAINTEXT://kafka:9092`(以及配套的监听变量)。 + +### `Invalid schema type for ByteArrayConverter: STRING` + +`ByteArrayFormat` 写入器只接受原始字节。要么把 worker 的 `value.converter` 换成 ByteArray converter,要么选择与当前 converter 匹配的 format class。 + +### `Class org.apache.kafka.connect.storage.ByteArrayConverter could not be found` + +Kafka 4.0 把该类移动到了 `org.apache.kafka.connect.converters.ByteArrayConverter`。请在 `worker.properties` 中使用新的包路径。 + +### 连接器已下载但插件未找到 + +Maven 上的裸连接器 JAR 不含依赖。请使用第 2 步的 Confluent Hub 压缩包(内含完整的 `lib/` 目录),并确保 `plugin.path` 指向包含连接器目录的上级目录。 + +## 下一步 + +- 在启用更多连接器前,先阅读 [S3 兼容性说明](/administration/protocols/s3)。 +- 使用[访问密钥管理](/security-compliance/iam/access-token)创建专用的生产凭证。 +- 按照 [Kafka Connect S3 sink 文档](https://docs.confluent.io/kafka-connect-s3/current/index.html)了解 Parquet 与 Avro 格式、按时间分区以及基于 IAM 的凭证链。 diff --git a/content/zh/developer/integration/big-data/meta.json b/content/zh/developer/integration/big-data/meta.json index f18ae39d..a4a40862 100644 --- a/content/zh/developer/integration/big-data/meta.json +++ b/content/zh/developer/integration/big-data/meta.json @@ -2,12 +2,15 @@ "title": "数据分析", "pages": [ "clickhouse", + "airflow", "duckdb", "doris", + "delta-lake", "flink", "hudi", "iceberg", "influxdb", + "kafka", "lakefs", "milvus", "mlflow", diff --git a/content/zh/developer/integration/devops/images/rustfs-opensearch-snapshot.png b/content/zh/developer/integration/devops/images/rustfs-opensearch-snapshot.png new file mode 100644 index 00000000..7f71a66b Binary files /dev/null and b/content/zh/developer/integration/devops/images/rustfs-opensearch-snapshot.png differ diff --git a/content/zh/developer/integration/devops/index.md b/content/zh/developer/integration/devops/index.md index a63e87bd..876949cc 100644 --- a/content/zh/developer/integration/devops/index.md +++ b/content/zh/developer/integration/devops/index.md @@ -8,6 +8,7 @@ description: "通过 S3 兼容的对象存储接口,将 DevOps 平台与基础 ## 平台与工具 - [Elasticsearch](./elasticsearch.md) +- [OpenSearch](./opensearch.md) - [Gitea](./gitea.md) - [Jenkins](./jenkins.md) - [Terraform](./terraform.md) diff --git a/content/zh/developer/integration/devops/meta.json b/content/zh/developer/integration/devops/meta.json index 67c84c6e..55eada6a 100644 --- a/content/zh/developer/integration/devops/meta.json +++ b/content/zh/developer/integration/devops/meta.json @@ -2,6 +2,7 @@ "title": "DevOps", "pages": [ "elasticsearch", + "opensearch", "gitea", "jenkins", "terraform" diff --git a/content/zh/developer/integration/devops/opensearch.md b/content/zh/developer/integration/devops/opensearch.md new file mode 100644 index 00000000..51ce035e --- /dev/null +++ b/content/zh/developer/integration/devops/opensearch.md @@ -0,0 +1,172 @@ +--- +title: "OpenSearch" +description: "用 repository-s3 插件把 OpenSearch 索引快照保存到 RustFS。" +--- + +本指南将源自 Elasticsearch 的开源搜索与分析套件 [OpenSearch](https://github.com/opensearch-project/OpenSearch) 通过 `repository-s3` 插件连接到 **RustFS**。你将注册一个以 RustFS 桶为后端的 S3 快照仓库,对索引打快照并还原。整个流程使用 `opensearchproject/opensearch:3.8.0` 与自带安装的 `repository-s3` 插件对 `rustfs/rustfs-x86-musl:v2.3.1` 验证通过。 + +你需要安装 Docker,或一个可以安装插件并修改配置的 OpenSearch 节点。本部署用于本地集成测试,不适用于生产环境。 + +## 架构 + +```mermaid +flowchart LR + Client["REST client"] --> OS["OpenSearch :9200"] + OS -->|"snapshot files"| RustFS["RustFS :9000"] + RustFS -->|"restore"| OS +``` + +`repository-s3` 插件把快照写成分片归档与元数据 blob 存入桶中。仓库注册是集群级配置,每个节点都需要该插件和相同的客户端配置。 + +## 1. 运行 OpenSearch + +以关闭安全插件和小堆内存启动单节点: + +```bash +docker run -d --name opensearch --network oo-rustfs_default -p 9200:9200 \ + -e discovery.type=single-node \ + -e OPENSEARCH_JAVA_OPTS="-Xms512m -Xmx512m" \ + -e DISABLE_SECURITY_PLUGIN=true \ + opensearchproject/opensearch:3.8.0 +``` + +当 `curl http://localhost:9200` 返回集群信息(需一到两分钟)即表示节点就绪。 + +## 2. 安装 repository-s3 插件 + +镜像没有预装 S3 仓库插件。安装后重启节点: + +```bash +docker exec opensearch bin/opensearch-plugin install --batch repository-s3 +docker restart opensearch +``` + +## 3. 配置 S3 客户端 + +凭证属于安全设置:必须放进 OpenSearch keystore,不能写在仓库请求或 `opensearch.yml` 里。创建 keystore 条目并替换全部连接占位符: + +```bash +docker exec opensearch sh -c \ + "printf '' | bin/opensearch-keystore create 2>/dev/null; \ + printf '' | bin/opensearch-keystore add -f -x s3.client.default.access_key; \ + printf '' | bin/opensearch-keystore add -f -x s3.client.default.secret_key" +``` + +把非敏感的客户端设置加入 `config/opensearch.yml`: + +```yaml title="opensearch.yml" +network.host: 0.0.0.0 +plugins.security.disabled: true +s3.client.default.endpoint: http://:9000 +s3.client.default.protocol: http +s3.client.default.path_style_access: "true" +``` + +再次重启节点,使其读取 keystore 与新配置: + +```bash +docker restart opensearch +``` + +趁节点启动时创建桶: + +```bash +rc mb rustfs/opensearch-snapshots +``` + +## 4. 注册仓库并打快照 + +创建测试索引和文档,然后注册仓库: + +```bash +curl -sX PUT http://localhost:9200/rustfs-demo -H "Content-Type: application/json" \ + -d '{"settings":{"number_of_shards":1}}' + +curl -sX PUT http://localhost:9200/rustfs-demo/_doc/1 -H "Content-Type: application/json" \ + -d '{"product":"rustfs","via":"opensearch-snapshot"}' + +curl -sX PUT "http://localhost:9200/_snapshot/rustfs-repo" -H "Content-Type: application/json" \ + -d '{"type":"s3","settings":{"bucket":"opensearch-snapshots","region":"us-east-1","server_side_encryption_type":"bucket_default"}}' +``` + +`server_side_encryption_type: bucket_default` 这个设置很关键:不设置时插件会请求 SSE-S3,而未配置服务端加密主密钥的自托管 RustFS 会拒绝它。 + +打快照并等待完成: + +```bash +curl -sX PUT "http://localhost:9200/_snapshot/rustfs-repo/snapshot-1?wait_for_completion=true" \ + -H "Content-Type: application/json" -d '{"indices":"rustfs-demo"}' +``` + +```text +{"snapshot":{"snapshot":"snapshot-1","state":"SUCCESS","indices":["rustfs-demo"],...}} +``` + +## 5. 验证对象并还原 + +列举桶: + +```bash +rc ls rustfs/opensearch-snapshots/ -r +``` + +```text +index-0 +index.latest +indices/5x1bwsWaSv2XINIbeoe-RQ/0/__GgxvoCw-TKuMMAYBq5Khag +indices/5x1bwsWaSv2XINIbeoe-RQ/0/snap-kRgFBiuyQIyPo3_-CMp7Hw.dat +meta-kRgFBiuyQIyPo3_-CMp7Hw.dat +snap-kRgFBiuyQIyPo3_-CMp7Hw.dat +``` + +删除索引并从快照还原: + +```bash +curl -sX DELETE http://localhost:9200/rustfs-demo +curl -sX POST "http://localhost:9200/_snapshot/rustfs-repo/snapshot-1/_restore?wait_for_completion=true" +curl -s http://localhost:9200/rustfs-demo/_doc/1 +``` + +```text +{"_index":"rustfs-demo","_id":"1","found":true,"_source":{"product":"rustfs","via":"opensearch-snapshot"}} +``` + +![存储在 RustFS 控制台中的 OpenSearch 快照](./images/rustfs-opensearch-snapshot.png) + +## 6. 停止或重置 + +保留桶内对象、仅拆除演示环境: + +```bash +docker rm -f opensearch +``` + +删除已存储的快照: + +```bash +rc rm rustfs/opensearch-snapshots/ --recursive --force +``` + +## 故障排查 + +### `Setting [access_key] is insecure, but property [allow_insecure_settings] is not set` + +仓库请求里的内联凭证会被拒绝。按第 3 步把凭证存入 keystore——在 OpenSearch 中 `access_key` 与 `secret_key` 是安全设置。 + +### `SSE-S3 requires RUSTFS_SSE_S3_MASTER_KEY ... (Status Code: 400)` + +插件默认用 SSE-S3 加密上传内容。按第 4 步以 `"server_side_encryption_type": "bucket_default"` 注册仓库,不发送加密头。 + +### 启动时报 `unknown setting [s3.client.default.access_key]` + +这些设置依赖 repository-s3 插件。如果节点带这些配置启动失败,说明该容器里没有插件——重复第 2 步(通过 `docker exec` 安装的插件在容器重建后会丢失)。 + +### 仓库验证报 `path is not accessible` + +节点无法访问桶:检查容器能否连通 `s3.client.default.endpoint`、`path_style_access` 是否为 `"true"`,以及 keystore 凭证是否已加载(凭证在启动时读取——添加后需要重启)。 + +## 下一步 + +- 在启用更多 OpenSearch 仓库前,先阅读 [S3 兼容性说明](/administration/protocols/s3)。 +- 使用[访问密钥管理](/security-compliance/iam/access-token)创建专用的生产凭证。 +- 按照 [OpenSearch 快照文档](https://docs.opensearch.org/docs/latest/tuning-your-cluster/availability-and-recovery/snapshots/index/)用 Snapshot Management(SM)策略自动执行快照。 diff --git a/content/zh/developer/integration/index.md b/content/zh/developer/integration/index.md index ff68281e..131694e6 100644 --- a/content/zh/developer/integration/index.md +++ b/content/zh/developer/integration/index.md @@ -7,14 +7,14 @@ description: "将 RustFS 与反向代理、备份工具、数据分析系统、A ## 集成类别 -- [反向代理](./reverse-proxy/index.md)涵盖 Nginx、Traefik、Caddy 和 HAProxy。 +- [反向代理](./reverse-proxy/index.md)涵盖 Nginx、Traefik、Caddy、HAProxy 和 Envoy。 - [备份](./backup/index.md)涵盖 Kopia、Longhorn、Restic 和 Velero。 -- [AI](./ai/index.md)涵盖 Ray 等 AI 平台。 -- [数据分析](./big-data/index.md)涵盖 ClickHouse、Doris、Hudi、Iceberg、lakeFS、Milvus、OpenDAL、Vitess 和 Zeppelin 等数据分析系统。 +- [AI](./ai/index.md)涵盖 Ray、vLLM 等 AI 平台。 +- [数据分析](./big-data/index.md)涵盖 Airflow、ClickHouse、Delta Lake、Doris、Hudi、Iceberg、Kafka、lakeFS、Milvus、OpenDAL、Vitess 和 Zeppelin 等数据分析系统。 - [云原生](./cloud-native/index.md)涵盖 Cortex 与 Flux。 -- [可观测性](./observability/index.md)涵盖 Fluentd、OpenObserve、OpenTelemetry、Thanos 和 Tempo 等遥测系统。 -- [其他](./others/index.md)涵盖社区驱动的 Python capo SDK。 +- [可观测性](./observability/index.md)涵盖 Fluentd、GreptimeDB、Loki、OpenObserve、OpenTelemetry、Tempo、Thanos 和 VictoriaMetrics 等遥测系统。 +- [其他](./others/index.md)涵盖 capo SDK、rclone、JuiceFS、Nextcloud 和 tusd 等工具。 - [镜像仓库](./registry/index.md)涵盖 Harbor。 -- [DevOps](./devops/index.md)涵盖 Elasticsearch、Gitea、Jenkins 和 Terraform。 +- [DevOps](./devops/index.md)涵盖 Elasticsearch、Gitea、Jenkins、OpenSearch 和 Terraform。 每篇指南都会说明配置集成系统时需要使用的 RustFS 端点和寻址要求。 \ No newline at end of file diff --git a/content/zh/developer/integration/observability/greptimedb.md b/content/zh/developer/integration/observability/greptimedb.md new file mode 100644 index 00000000..a3b9f5ae --- /dev/null +++ b/content/zh/developer/integration/observability/greptimedb.md @@ -0,0 +1,129 @@ +--- +title: "GreptimeDB" +description: "以 RustFS 作为 GreptimeDB 的 S3 兼容对象存储后端。" +--- + +本指南将云原生开源时序数据库 [GreptimeDB](https://github.com/GreptimeTeam/greptimedb) 连接到 **RustFS** 作为其对象存储后端。你将启动一个 `[storage]` 配置指向 RustFS 桶的 standalone 实例,通过 SQL API 写入时序数据,并确认桶内的 Parquet 文件与 manifest。整个流程使用 `greptime/greptimedb`(main,commit `179ff8e5`)对 `rustfs/rustfs-x86-musl:v2.3.1` 验证通过。 + +你需要安装 Docker,或本地 GreptimeDB 二进制。本部署用于本地集成测试,不适用于生产环境。 + +## 架构 + +```mermaid +flowchart LR + SQL["SQL / Prometheus API"] --> DB["GreptimeDB"] + DB -->|"SST + manifests"| RustFS["RustFS :9000"] +``` + +GreptimeDB 把 WAL 和最新数据保留在本地,随后将 SSTable(Parquet)与表 manifest 持久化到对象存储。把存储后端指向 RustFS 后,该桶就是所有表数据的持久化归属地。 + +## 1. 配置存储后端 + +创建桶和一个带 S3 存储段的配置文件,替换全部连接占位符: + +```toml title="greptimedb.toml" +[storage] +type = "S3" +bucket = "" +root = "greptimedb" +access_key_id = "" +secret_access_key = "" +endpoint = "http://:9000" +region = "us-east-1" +``` + +GreptimeDB 对自定义端点默认使用路径风格请求;虚拟主机风格需要显式开启 `enable_virtual_host_style`,因此对接 RustFS 无需额外参数。 + +## 2. 运行 GreptimeDB + +以配置文件启动 standalone 实例: + +```bash +docker run -d --name greptimedb --network oo-rustfs_default -p 4000:4000 -p 4002:4002 \ + -v "$PWD/greptimedb.toml":/etc/greptimedb/greptimedb.toml:ro \ + greptime/greptimedb:latest standalone start \ + --http-addr 0.0.0.0:4000 \ + --mysql-addr 0.0.0.0:4002 \ + --config-file /etc/greptimedb/greptimedb.toml +``` + +`4000` 端口提供 HTTP SQL 端点,`4002` 提供 MySQL 协议。 + +## 3. 写入并查询时序数据 + +建表、插入、回读。HTTP SQL 端点接受表单编码请求: + +```bash +curl -s -X POST "http://localhost:4000/v1/sql" \ + --data-urlencode "sql=CREATE TABLE rustfs_demo (host STRING, cpu DOUBLE, mem DOUBLE, ts TIMESTAMP TIME INDEX)" + +curl -s -X POST "http://localhost:4000/v1/sql" \ + --data-urlencode "sql=INSERT INTO rustfs_demo VALUES (\"node-1\", 0.31, 0.62, 1790681000000), (\"node-1\", 0.35, 0.63, 1790681060000), (\"node-2\", 0.51, 0.71, 1790681000000)" +``` + +```text +{"output":[{"affectedrows":3}],"execution_time_ms":2} +``` + +查询回读: + +```bash +curl -s -X POST "http://localhost:4000/v1/sql" \ + --data-urlencode "sql=SELECT * FROM rustfs_demo ORDER BY ts" +``` + +```text +{"output":[{"records":{"rows":[["node-2",0.51,0.71,1790681000000],["node-1",0.35,0.63,1790681060000]],"total_rows":2}}]} +``` + +## 4. 验证 RustFS 中的对象 + +列举桶——memtable 刷盘后,桶内会出现 Parquet SSTable 和 JSON manifest: + +```bash +rc ls rustfs// -r +``` + +```text +greptimedb/data/greptime/public/1024/1024_0000000000/manifest/00000000000000000000.json +greptimedb/data/greptime/greptime_private/1025/1025_0000000000/b11e8b25-5763-4f05-bcab-b6ee0a756a69.parquet +greptimedb/data/greptime/greptime_private/1025/1025_0000000000/manifest/00000000000000000001.json +``` + +每个数据库在 `data/` 下有自己的目录,按 region 存放的 `manifest/*.json` 描述了查询时读回的 SSTable。 + +![存储在 RustFS 控制台中的 GreptimeDB 数据](./images/rustfs-greptimedb-data.png) + +## 5. 停止或重置 + +保留桶内对象、仅拆除演示环境: + +```bash +docker rm -f greptimedb +``` + +删除已存储的数据: + +```bash +rc rm rustfs// --recursive --force +``` + +## 故障排查 + +### `Form requests must have Content-Type: application/x-www-form-urlencoded` + +`/v1/sql` HTTP 端点只接受表单编码请求体。请用 `curl --data-urlencode "sql=..."`(或 `application/x-www-form-urlencoded`)传 SQL,不要用 JSON 请求体。 + +### 桶一直是空的 + +GreptimeDB 异步地把 memtable 刷到对象存储。多写几条并等几秒(或手动触发 flush),再列举桶。 + +### 启动时报 S3 错误 + +确认 `endpoint` 带协议前缀、桶已创建,且 `access_key_id`/`secret_access_key` 与 RustFS 访问密钥匹配。`root` 可选,但建议保留以便所有表目录都在已知前缀下。 + +## 下一步 + +- 在启用更多 GreptimeDB 存储选项前,先阅读 [S3 兼容性说明](/administration/protocols/s3)。 +- 使用[访问密钥管理](/security-compliance/iam/access-token)创建专用的生产凭证。 +- 按照 [GreptimeDB 配置参考](https://docs.greptime.com/operational-guide/configure/configure-datanode/)为生产负载调整 flush 周期与缓存层。 diff --git a/content/zh/developer/integration/observability/images/rustfs-greptimedb-data.png b/content/zh/developer/integration/observability/images/rustfs-greptimedb-data.png new file mode 100644 index 00000000..8f43aa40 Binary files /dev/null and b/content/zh/developer/integration/observability/images/rustfs-greptimedb-data.png differ diff --git a/content/zh/developer/integration/observability/images/rustfs-vm-backups.png b/content/zh/developer/integration/observability/images/rustfs-vm-backups.png new file mode 100644 index 00000000..a47232c5 Binary files /dev/null and b/content/zh/developer/integration/observability/images/rustfs-vm-backups.png differ diff --git a/content/zh/developer/integration/observability/index.md b/content/zh/developer/integration/observability/index.md index 8b6a450b..b3286070 100644 --- a/content/zh/developer/integration/observability/index.md +++ b/content/zh/developer/integration/observability/index.md @@ -8,10 +8,12 @@ description: "通过 S3 兼容对象存储接口,将可观测性平台连接 ## 平台 - [Fluentd](./fluentd.md) +- [GreptimeDB](./greptimedb.md) - [OpenObserve](./openobserve.md) - [OpenTelemetry](./opentelemetry.md) - [Loki](./loki.md) - [Tempo](./tempo.md) - [Thanos](./thanos.md) +- [VictoriaMetrics](./victoriametrics.md) 请使用专用的存储桶保存遥测数据,并为凭证仅授予所需桶操作的权限。 diff --git a/content/zh/developer/integration/observability/meta.json b/content/zh/developer/integration/observability/meta.json index 0acec73a..7ef6a1a6 100644 --- a/content/zh/developer/integration/observability/meta.json +++ b/content/zh/developer/integration/observability/meta.json @@ -2,10 +2,12 @@ "title": "可观测性", "pages": [ "fluentd", + "greptimedb", "loki", "openobserve", "opentelemetry", "tempo", - "thanos" + "thanos", + "victoriametrics" ] } diff --git a/content/zh/developer/integration/observability/victoriametrics.md b/content/zh/developer/integration/observability/victoriametrics.md new file mode 100644 index 00000000..3ded99a7 --- /dev/null +++ b/content/zh/developer/integration/observability/victoriametrics.md @@ -0,0 +1,157 @@ +--- +title: "VictoriaMetrics" +description: "用 vmbackup 把 VictoriaMetrics 快照备份到 RustFS。" +--- + +本指南将 Prometheus 兼容时序数据库 [VictoriaMetrics](https://github.com/VictoriaMetrics/VictoriaMetrics) 通过 `vmbackup` 与 `vmrestore` 连接到 **RustFS**。你将运行单节点实例、导入指标、创建即时快照、备份到 RustFS 桶,并把数据还原到全新目录。整个流程使用 `victoria-metrics`、`vmbackup` 与 `vmrestore` v1.x 镜像对 `rustfs/rustfs-x86-musl:v2.3.1` 验证通过。 + +你需要安装 Docker。本部署用于本地集成测试,不适用于生产环境。 + +## 架构 + +```mermaid +flowchart LR + Import["Prometheus import API"] --> VM["VictoriaMetrics :8428"] + VM -->|"instant snapshot"| Backup["vmbackup"] + Backup -->|"S3 API"| RustFS["RustFS :9000"] + RustFS -->|"restore"| Restore["vmrestore"] +``` + +`vmbackup` 把存储目录在某一时间点的一致性快照上传到任意 S3 兼容端点。`vmrestore` 执行相反过程,产出一个 VictoriaMetrics 实例可通过 -storageDataPath 直接打开的数据目录。 + +## 1. 运行 VictoriaMetrics + +创建桶并启动单节点实例: + +```bash +rc mb rustfs/vm-backups + +docker run -d --name vm --network oo-rustfs_default -p 8428:8428 \ + -v vm-data:/storage \ + victoriametrics/victoria-metrics:latest \ + -storageDataPath=/storage -retentionPeriod=100y +``` + +## 2. 导入指标 + +通过 Prometheus import API 写入两条样本: + +```bash +echo "vm_demo_metric 123" | curl -s --data-binary @- http://localhost:8428/api/v1/import/prometheus +echo "vm_demo_metric 456" | curl -s --data-binary @- http://localhost:8428/api/v1/import/prometheus +``` + +该端点返回 `204 No Content`。确认数据可查: + +```bash +curl -s "http://localhost:8428/api/v1/export?match[]=vm_demo_metric" +``` + +```text +{"metric":{"__name__":"vm_demo_metric"},"values":[123,456],"timestamps":[1790680894604,1790680894619]} +``` + +## 3. 创建快照 + +向 VictoriaMetrics 请求一致性快照: + +```bash +curl -s http://localhost:8428/snapshot/create +``` + +```text +{"status":"ok","snapshot":"20260929112134-18D9C6C76704F913"} +``` + +## 4. 把快照备份到 RustFS + +在同一个存储卷上运行 `vmbackup`,替换凭证占位符。快照名来自第 3 步: + +```bash +docker run --rm --network oo-rustfs_default \ + -e AWS_ACCESS_KEY_ID= \ + -e AWS_SECRET_ACCESS_KEY= \ + --volumes-from vm \ + victoriametrics/vmbackup:latest \ + -storageDataPath=/storage \ + -snapshotName=20260929112134-18D9C6C76704F913 \ + -dst=s3://vm-backups/demo \ + -customS3Endpoint=http://:9000 +``` + +```text +backup ... to S3{bucket: "vm-backups", dir: "demo/"} is complete; uploaded 760 bytes +``` + +`-customS3Endpoint` 把 AWS SDK 重定向到 RustFS;自定义端点自动使用路径风格寻址。在没有 EC2 元数据服务的主机上,建议同时设置 `AWS_EC2_METADATA_DISABLED=true`,避免凭证探测带来的延迟。 + +## 5. 验证与还原 + +列举桶内前缀: + +```bash +rc ls rustfs/vm-backups/demo/ +``` + +```text +backup_complete.ignore +backup_metadata.ignore +data/ +metadata/ +``` + +`backup_complete.ignore` 标记一次完整备份。把它还原到全新目录: + +```bash +docker run --rm --network oo-rustfs_default \ + -e AWS_ACCESS_KEY_ID= \ + -e AWS_SECRET_ACCESS_KEY= \ + -v /opt/vm-restore:/restore \ + victoriametrics/vmrestore:latest \ + -src=s3://vm-backups/demo \ + -storageDataPath=/restore \ + -customS3Endpoint=http://:9000 +``` + +```text +restored 760 bytes from backup in 0.055 seconds +``` + +还原出的目录包含 `data/`、`metadata/` 和锁文件——正是 VictoriaMetrics 实例在 `-storageDataPath` 处期望的内容。 + +![存储在 RustFS 控制台中的 VictoriaMetrics 备份](./images/rustfs-vm-backups.png) + +## 6. 停止或重置 + +保留桶内对象、仅拆除演示环境: + +```bash +docker rm -f vm +docker volume rm vm-data +``` + +删除已存储的备份: + +```bash +rc rm rustfs/vm-backups/ --recursive --force +``` + +## 故障排查 + +### `vmbackup` 启动卡住或找不到凭证 + +缺少环境变量凭证时,AWS SDK 会探测 EC2 元数据服务。在没有 IMDS 的机器上,请在密钥变量之外导出 `AWS_EC2_METADATA_DISABLED=true`。 + +### 每次运行都重新上传全部备份分块 + +`vmbackup` 通过比对本地与远端文件的哈希做增量备份。还原到新目录后再从那里执行 `vmbackup` 会全量重传;增量运行请保留原始数据目录。 + +### 导入后立即查询不到数据 + +导入是异步接收的,繁忙的单节点上即时查询端点可能有短暂延迟。创建快照前先用 `/api/v1/export` 验证(或等几秒)。 + +## 下一步 + +- 在启用更多 VictoriaMetrics 组件前,先阅读 [S3 兼容性说明](/administration/protocols/s3)。 +- 使用[访问密钥管理](/security-compliance/iam/access-token)创建专用的生产凭证。 +- 按照 [vmbackup 文档](https://docs.victoriametrics.com/vmbackup/)安排备份计划并清理旧快照。 diff --git a/content/zh/developer/integration/others/images/rustfs-juicefs-chunks.png b/content/zh/developer/integration/others/images/rustfs-juicefs-chunks.png new file mode 100644 index 00000000..5f1bd211 Binary files /dev/null and b/content/zh/developer/integration/others/images/rustfs-juicefs-chunks.png differ diff --git a/content/zh/developer/integration/others/images/rustfs-nextcloud-file.png b/content/zh/developer/integration/others/images/rustfs-nextcloud-file.png new file mode 100644 index 00000000..2015f14a Binary files /dev/null and b/content/zh/developer/integration/others/images/rustfs-nextcloud-file.png differ diff --git a/content/zh/developer/integration/others/images/rustfs-rclone-sync.png b/content/zh/developer/integration/others/images/rustfs-rclone-sync.png new file mode 100644 index 00000000..bed1114c Binary files /dev/null and b/content/zh/developer/integration/others/images/rustfs-rclone-sync.png differ diff --git a/content/zh/developer/integration/others/images/rustfs-tus-uploads.png b/content/zh/developer/integration/others/images/rustfs-tus-uploads.png new file mode 100644 index 00000000..08f8aabe Binary files /dev/null and b/content/zh/developer/integration/others/images/rustfs-tus-uploads.png differ diff --git a/content/zh/developer/integration/others/index.md b/content/zh/developer/integration/others/index.md index 94b85318..eafce2e9 100644 --- a/content/zh/developer/integration/others/index.md +++ b/content/zh/developer/integration/others/index.md @@ -8,3 +8,7 @@ description: "其他 RustFS 集成,目前为社区驱动的 Python capo SDK。 ## 指南 - [capo (Python)](./capo.md) —— 使用同步或异步客户端将社区驱动的 capo SDK 连接到 RustFS。 +rclone](./rclone.md) — sync, mount, and serve RustFS buckets from the command line. +- [tusd](./tusd.md) — receive resumable uploads into a RustFS bucket over the tus protocol. +- [JuiceFS](./juicefs.md) — mount a POSIX filesystem backed by a RustFS bucket. +- [Nextcloud](./nextcloud.md) — use RustFS as S3 external storage for Nextcloud files. diff --git a/content/zh/developer/integration/others/juicefs.md b/content/zh/developer/integration/others/juicefs.md new file mode 100644 index 00000000..9f54e0ef --- /dev/null +++ b/content/zh/developer/integration/others/juicefs.md @@ -0,0 +1,132 @@ +--- +title: "JuiceFS" +description: "用 JuiceFS 的 S3 对象存储后端,在 RustFS 上构建 POSIX 文件系统。" +--- + +本指南将云原生分布式 POSIX 文件系统 [JuiceFS](https://github.com/juicedata/juicefs) 连接到 **RustFS** 作为其对象存储后端。你将格式化一个数据块存放在 RustFS 桶中的卷,挂载到本地,并通过挂载点读写文件。整个流程使用 `juicefs v1.3.1`(社区版,SQLite 元数据引擎)对 `rustfs/rustfs-x86-musl:v2.3.1` 验证通过。 + +你需要 JuiceFS 二进制、一个元数据引擎以及 FUSE(Linux 上为 `fuse3`)。本部署用于本地集成测试,不适用于生产环境。 + +## 架构 + +```mermaid +flowchart LR + Mount["/mnt/jfs"] -->|"POSIX"| JuiceFS["JuiceFS client"] + JuiceFS -->|"metadata"| Meta["SQLite / Redis"] + JuiceFS -->|"data chunks"| RustFS["RustFS :9000"] +``` + +JuiceFS 把每个文件切成块,以对象的形式存放在桶内 `chunks/` 目录下,元数据引擎负责记录文件名、inode 与布局。文件系统表现得像本地磁盘,但不在本地保留任何数据。 + +## 1. 格式化卷 + +创建存储桶并格式化一个以 RustFS 为后端的 JuiceFS 卷,替换全部连接占位符。桶 URL 中内嵌的端点决定了使用路径风格寻址: + +```bash +rc mb rustfs/jfs-demo + +juicefs format \ + --storage s3 \ + --bucket http://:9000/jfs-demo \ + --access-key \ + --secret-key \ + sqlite3:///opt/juicefs/jfs.db \ + rustfs-jfs +``` + +```text +Data use s3://:9000/jfs-demo/rustfs-jfs/ + OK, rustfs-jfs is ready +``` + +`sqlite3:///opt/juicefs/jfs.db` 是本次测试的元数据引擎。生产环境请改用 Redis、MySQL 或 PostgreSQL,以便多个客户端挂载同一个卷。 + +## 2. 挂载卷 + +使用相同的元数据 URL 挂载文件系统: + +```bash +mkdir -p /mnt/jfs +juicefs mount -d sqlite3:///opt/juicefs/jfs.db /mnt/jfs +``` + +```text +OK, rustfs-jfs is ready at /mnt/jfs +``` + +`-d` 表示后台运行。该卷现在就是一个 POSIX 文件系统。 + +## 3. 读写文件 + +像使用普通目录一样使用挂载点: + +```bash +echo "hello rustfs jfs" > /mnt/jfs/hello.txt +dd if=/dev/urandom of=/mnt/jfs/blob.bin bs=1M count=3 +mkdir -p /mnt/jfs/dir1 && echo nested > /mnt/jfs/dir1/nested.txt +cat /mnt/jfs/hello.txt +``` + +```text +hello rustfs jfs +``` + +用 `juicefs info` 查看卷的状态: + +```text +/mnt/jfs : + inode: 1 + files: 2 + dirs: 2 + length: 3.00 MiB +``` + +## 4. 验证 RustFS 中的数据块 + +列举桶内前缀: + +```bash +rc ls rustfs/jfs-demo/ -r +``` + +每个文件都被切成了按内容寻址的块对象: + +```text +rustfs-jfs/chunks/0/0/1_0_17 +rustfs-jfs/chunks/0/0/3_0_3145728 +rustfs-jfs/chunks/0/0/4_0_7 +``` + +块名编码了 inode、块序号和大小——例如 `3_0_3145728` 就是第 3 步写入的 3 MiB 文件。 + +![存储在 RustFS 控制台中的 JuiceFS 数据块](./images/rustfs-juicefs-chunks.png) + +## 5. 停止或重置 + +卸载卷,然后可选地销毁卷元数据和桶内数据: + +```bash +juicefs umount /mnt/jfs +juicefs destroy --force sqlite3:///opt/juicefs/jfs.db rustfs-jfs +rc rm rustfs/jfs-demo/ --recursive --force +``` + +## 故障排查 + +### `unknown option: --daemon` + +后台运行的 flag 是短横线一个:`juicefs mount -d`。不带它运行则挂载保持在前台(便于调试)。 + +### `fusermount3: mount failed: Permission denied` + +挂载需要 FUSE 设备。在容器内运行时加 `--device /dev/fuse --cap-add SYS_ADMIN`(或 `--privileged`);在宿主机上安装 `fuse3` 并确认 `/dev/fuse` 存在。 + +### 挂载卡住或访问不到存储 + +客户端需要同时访问元数据引擎和桶端点。由于桶 URL 内嵌了端点,格式化之前先在挂载主机上用 `curl` 验证连通性。 + +## 下一步 + +- 在启用更多 JuiceFS 存储后端前,先阅读 [S3 兼容性说明](/administration/protocols/s3)。 +- 使用[访问密钥管理](/security-compliance/iam/access-token)创建专用的生产凭证。 +- 按照 [JuiceFS 文档](https://juicefs.com/docs/community/quick_start_guide/)把元数据引擎切换到 Redis,并从多个客户端挂载同一卷。 diff --git a/content/zh/developer/integration/others/meta.json b/content/zh/developer/integration/others/meta.json index 588b5c2d..3b601ed8 100644 --- a/content/zh/developer/integration/others/meta.json +++ b/content/zh/developer/integration/others/meta.json @@ -1,6 +1,10 @@ { "title": "其他", "pages": [ - "capo" + "capo", + "juicefs", + "nextcloud", + "rclone", + "tusd" ] } diff --git a/content/zh/developer/integration/others/nextcloud.md b/content/zh/developer/integration/others/nextcloud.md new file mode 100644 index 00000000..792842e0 --- /dev/null +++ b/content/zh/developer/integration/others/nextcloud.md @@ -0,0 +1,145 @@ +--- +title: "Nextcloud" +description: "将 RustFS 用作 Nextcloud 的 S3 外部存储。" +--- + +本指南将自托管内容协作平台 [Nextcloud](https://github.com/nextcloud/server) 通过其 External Storage 应用的 S3 后端连接到 **RustFS**。你将启用 `files_external`,把一个 RustFS 桶挂载到所有用户的文件视图中,并通过 WebDAV 上传一个直接落入桶内的文件。整个流程使用 `nextcloud:32.0.15`(SQLite,单容器)对 `rustfs/rustfs-x86-musl:v2.3.1` 验证通过。 + +你需要安装 Docker,或一个可以使用 `occ` 的现有 Nextcloud 实例。本部署用于本地集成测试,不适用于生产环境。 + +## 架构 + +```mermaid +flowchart LR + User["Browser / WebDAV"] --> Nextcloud["Nextcloud"] + Nextcloud -->|"files_external (S3)"| RustFS["RustFS :9000"] +``` + +Nextcloud 把挂载点上的文件操作代理到 S3 后端。对象按其在挂载内的相对路径存放,因此桶里的名称与用户看到的名称一致。 + +## 1. 安装 Nextcloud + +带管理员账号运行 Nextcloud,并替换全部连接占位符。SQLite 让测试环境自成一体;生产环境请使用 MariaDB 或 PostgreSQL: + +```bash +docker run -d --name nextcloud --network oo-rustfs_default -p 8080:80 \ + -e NEXTCLOUD_ADMIN_USER= \ + -e NEXTCLOUD_ADMIN_PASSWORD= \ + nextcloud:32.0.15 +``` + +如果启动后网页仍显示安装向导,手动完成安装: + +```bash +docker exec -u www-data nextcloud php occ maintenance:install \ + --admin-user --admin-password +``` + +## 2. 启用 External Storage 应用 + +`files_external` 应用随 Nextcloud 一起分发但默认未启用,其 `occ` 命令只有在应用启用后才存在: + +```bash +docker exec -u www-data nextcloud php occ app:enable files_external +``` + +```text +files_external 1.24.1 enabled +``` + +## 3. 挂载 RustFS 桶 + +以后端类型 `amazons3`、认证后端 `amazons3::accesskey` 创建外部存储,并替换全部连接占位符: + +```bash +docker exec -u www-data nextcloud php occ files_external:create \ + /rustfs amazons3 amazons3::accesskey \ + --user \ + --config bucket= \ + --config hostname= \ + --config port=9000 \ + --config use_ssl=false \ + --config use_path_style=true \ + --config key= \ + --config secret= +``` + +```text +Storage created with id 1 +``` + +挂载点 `/rustfs` 会出现在该用户的文件视图中。非 AWS 端点必须设置 `use_path_style=true`。使用前先检查连接: + +```bash +docker exec -u www-data nextcloud php occ files_external:verify 1 +``` + +```text + - status: ok + - code: 0 +``` + +## 4. 上传文件并验证 + +通过 WebDAV 端点上传,写入会经由外部存储落盘: + +```bash +echo "nextcloud writes to rustfs" > /tmp/nc-demo.txt + +curl -u : \ + -T /tmp/nc-demo.txt \ + http://localhost:8080/remote.php/dav/files//rustfs/nc-demo.txt \ + -o /dev/null -w "%{http_code}\n" +``` + +```text +201 +``` + +回读该文件,然后确认 RustFS 中的对象: + +```bash +rc ls rustfs// -r +``` + +```text +[2026-09-29 13:42:48] 27 B nc-demo.txt +``` + +对象的键与挂载内的路径一致,因此通过 Nextcloud 上传的文件也可以被任何 S3 客户端直接读取。 + +![存储在 RustFS 控制台中的 Nextcloud 文件](./images/rustfs-nextcloud-file.png) + +## 5. 停止或重置 + +保留桶内对象、仅删除挂载: + +```bash +docker exec -u www-data nextcloud php occ files_external:delete 1 +``` + +删除桶内内容: + +```bash +rc rm rustfs// --recursive --force +``` + +## 故障排查 + +### `There are no commands defined in the "files_external" namespace` + +应用尚未启用。先执行 `occ app:enable files_external`,之后 `occ files_external:*` 命令才会注册。 + +### `Not enough arguments (missing: "authentication_backend")` + +`files_external:create` 需要分别传入存储后端与认证后端两个参数:`amazons3 amazons3::accesskey`。后端标识可通过 `occ files_external:backends` 查看。 + +### 挂载可见但为空,或上传失败 + +确认 Nextcloud 容器可以访问 `hostname`(使用容器网络名称,不要用 `localhost`)、`use_path_style` 为 `true`,且桶已存在。`occ files_external:verify ` 会给出具体的连接错误。 + +## 下一步 + +- 在启用更多外部存储后端前,先阅读 [S3 兼容性说明](/administration/protocols/s3)。 +- 使用[访问密钥管理](/security-compliance/iam/access-token)创建专用的生产凭证。 +- 按照 [Nextcloud 外部存储文档](https://docs.nextcloud.com/server/latest/admin_manual/configuration_files/external_storage_configuration_gui.html)把挂载共享给群组并启用版本管理。 diff --git a/content/zh/developer/integration/others/rclone.md b/content/zh/developer/integration/others/rclone.md new file mode 100644 index 00000000..ec78c47c --- /dev/null +++ b/content/zh/developer/integration/others/rclone.md @@ -0,0 +1,144 @@ +--- +title: "rclone" +description: "通过 rclone 的 S3 兼容 API 对 RustFS 存储桶执行 sync、mount 与 serve。" +--- + +本指南将对象存储领域的命令行"瑞士军刀"[rclone](https://github.com/rclone/rclone) 通过其 S3 后端连接到 **RustFS**。你将为 RustFS 配置一个 S3 remote,复制并同步文件、回读对象、用 `rclone serve` 把桶发布为 HTTP 服务,并用 `rclone mount` 把桶挂载为本地文件系统。整个流程使用 `rclone v1.75.1` 对 `rustfs/rustfs-x86-musl:v2.3.1` 验证通过。 + +你需要安装 Docker,或本地 rclone 二进制。本部署用于本地集成测试,不适用于生产环境。 + +## 架构 + +```mermaid +flowchart LR + Files["Local files"] -->|"copy / sync"| Remote["rclone S3 remote"] + Remote -->|"S3 API"| RustFS["RustFS :9000"] + RustFS -->|"mount / serve"| Client["FUSE mount / HTTP clients"] +``` + +一个 remote 定义即可驱动所有 rclone 命令:数据传输、挂载和发布都复用同一个 S3 连接。 + +## 1. 配置 remote + +创建 rclone 配置文件,为 RustFS 定义一个 S3 remote,并替换全部连接占位符。`Other` provider 会关闭 AWS 特有行为,自定义端点自动使用路径风格寻址: + +```ini title="rclone.conf" +[rustfs] +type = s3 +provider = Other +access_key_id = +secret_access_key = +endpoint = http://:9000 +region = us-east-1 +``` + +## 2. 复制并读取对象 + +创建存储桶并用 `rclone copy` 上传目录: + +```bash +rc mb rustfs/rclone-demo +rclone copy /data rustfs:rclone-demo/seed +``` + +列举并回读: + +```bash +rclone ls rustfs:rclone-demo/seed +rclone cat rustfs:rclone-demo/seed/hello.txt +``` + +```text + 3145728 blob.bin + 18 hello.txt +hello from rclone +``` + +`rclone lsd rustfs:` 可以列出端点上的所有存储桶。 + +## 3. 同步目录 + +`rclone sync` 让目标与源完全一致,包括删除。删除一个本地文件后同步: + +```bash +rm /data/hello.txt +rclone sync /data rustfs:rclone-demo/seed +rclone lsf rustfs:rclone-demo/seed +``` + +```text +blob.bin +``` + +`hello.txt` 从桶中消失。首次执行可先加 `--dry-run` 预览变更,不会触碰桶内数据。 + +## 4. 把桶发布为 HTTP 服务 + +把桶内容发布为 HTTP 文件服务: + +```bash +rclone serve http --addr 0.0.0.0:8080 rustfs:rclone-demo/seed +``` + +任何 HTTP 客户端都能下载对象: + +```bash +curl -s http://localhost:8080/blob.bin -o /dev/null -w "%{http_code} %{size_download} bytes\n" +``` + +```text +200 3145728 bytes +``` + +同一个 remote 还支持 `rclone serve` 的 WebDAV、SFTP 与 S3 端点。 + +## 5. 把桶挂载为文件系统 + +在有 FUSE 的环境中,把桶挂载到本地并像目录一样使用: + +```bash +rclone mount rustfs:rclone-demo /mnt/rclone --daemon +ls /mnt/rclone/seed +echo test > /mnt/rclone/write-test.txt +cat /mnt/rclone/write-test.txt +``` + +通过挂载点写入的文件会以普通对象的形式出现在 RustFS 中: + +```bash +rc ls rustfs/rclone-demo/ -r +``` + +```text +[2026-09-29 11:12:41] 5 B write-test.txt +``` + +使用完毕后执行 `fusermount -u /mnt/rclone` 卸载。 + +## 6. 停止或重置 + +rclone 不持有任何服务端状态。删除演示数据: + +```bash +rclone purge rustfs:rclone-demo +``` + +## 故障排查 + +### `Access Denied` 或列表为空 + +确认 `endpoint` 带协议前缀,密钥对与 RustFS 访问密钥匹配。虽然 RustFS 会忽略 `region`,但 S3 签名需要该值;保持 `us-east-1` 即可。 + +### 挂载报 `fusermount3: mount failed: Permission denied` + +挂载需要 FUSE 设备和提升的权限。在容器内运行时加 `--device /dev/fuse --cap-add SYS_ADMIN`——若挂载助手仍然失败,改用 `--privileged`。在宿主机上则确认已安装 `fuse3` 且 `/dev/fuse` 存在。 + +### 同步没有删除目标端的文件 + +`rclone copy` 从不删除。只有 `rclone sync`(或 `rclone delete`)会移除目标对象,且两者都可用 `--dry-run` 预览。 + +## 下一步 + +- 在启用更多 rclone 后端前,先阅读 [S3 兼容性说明](/administration/protocols/s3)。 +- 使用[访问密钥管理](/security-compliance/iam/access-token)创建专用的生产凭证。 +- 按照 [rclone S3 文档](https://rclone.org/s3/)了解 `--transfers`、带宽限制与 crypt 覆盖层等参数。 diff --git a/content/zh/developer/integration/others/tusd.md b/content/zh/developer/integration/others/tusd.md new file mode 100644 index 00000000..05067e15 --- /dev/null +++ b/content/zh/developer/integration/others/tusd.md @@ -0,0 +1,170 @@ +--- +title: "tusd" +description: "用 tusd 服务器的 S3 后端,把断点续传上传接收到 RustFS。" +--- + +本指南将 tus 断点续传协议的官方参考实现 [tusd](https://github.com/tus/tusd) 连接到 **RustFS** 作为其 S3 存储后端。你将针对 RustFS 桶运行 tusd,用 tus 协议创建上传、分两片发送文件(中间故意中断)、从服务器上报的偏移量续传,并验证桶内组装好的对象。整个流程使用 `tusproject/tusd:v2.10.1` 对 `rustfs/rustfs-x86-musl:v2.3.1` 验证通过。 + +你需要安装 Docker,或本地 tusd 二进制。本部署用于本地集成测试,不适用于生产环境。 + +## 架构 + +```mermaid +flowchart LR + Client["tus client"] -->|"POST / PATCH / HEAD"| tusd["tusd :8080"] + tusd -->|"multipart upload"| RustFS["RustFS :9000"] +``` + +tusd 把每个进行中的上传以 S3 分片上传的形式存入桶内。断线后的客户端用 `HEAD` 查询已提交的偏移量并从断点继续——已收到的数据永远不会重复传输。 + +## 1. 运行 tusd + +创建桶并以 S3 后端启动 tusd,替换全部连接占位符。虽然 RustFS 会忽略 region,但 AWS 环境必须设置它: + +```bash +rc mb rustfs/tus-uploads + +docker run -d --name tusd --network oo-rustfs_default -p 8080:8080 \ + -e AWS_ACCESS_KEY_ID= \ + -e AWS_SECRET_ACCESS_KEY= \ + -e AWS_REGION=us-east-1 \ + tusproject/tusd:latest \ + -s3-bucket tus-uploads \ + -s3-endpoint http://:9000 +``` + +检查服务健康: + +```bash +curl -s -o /dev/null -w "%{http_code}\n" http://localhost:8080/health +``` + +```text +200 +``` + +## 2. 创建上传 + +创建一个 6 MiB 的上传并读取 `Location` 头: + +```bash +curl -s -D - -o /dev/null -X POST http://localhost:8080/files/ \ + -H "Upload-Length: 6291456" -H "Tus-Resumable: 1.0.0" \ + | grep -i "^Location:" +``` + +```text +Location: http://localhost:8080/files/b7338250daa9...+NGVmNDRhZjEt... +``` + +上传 URL 包含文件 ID 和消息认证标签。再次发送前要剥掉协议与主机部分(服务器回显收到的 `Host`,它未必是下一个客户端可达的地址)。 + +## 3. 分片上传并模拟中断 + +先发前 2.5 MB,然后停下——这正是移动客户端断网的位置: + +```bash +head -c 2500000 demo.bin > part1.bin + +curl -s -o /dev/null -w "%{http_code}\n" -X PATCH "http://localhost:8080${LOC}" \ + -H "Upload-Offset: 0" -H "Tus-Resumable: 1.0.0" \ + -H "Content-Type: application/offset+octet-stream" \ + --data-binary @part1.bin +``` + +```text +204 +``` + +向服务器查询它实际收到了多少——这就是断点续传的核心: + +```bash +curl -s -X HEAD "http://localhost:8080${LOC}" \ + -H "Tus-Resumable: 1.0.0" -D - -o /dev/null | grep -i upload-offset +``` + +```text +Upload-Offset: 2500000 +``` + +## 4. 续传并完成 + +从偏移量 2500000 继续发送剩余字节: + +```bash +tail -c 3791456 demo.bin > part2.bin + +curl -s -o /dev/null -w "%{http_code}\n" -X PATCH "http://localhost:8080${LOC}" \ + -H "Upload-Offset: 2500000" -H "Tus-Resumable: 1.0.0" \ + -H "Content-Type: application/offset+octet-stream" \ + --data-binary @part2.bin +``` + +```text +204 +``` + +通过 tusd 下载完成的文件,并与源文件比对校验和: + +```bash +curl -s -o download.bin "http://localhost:8080${LOC}" +sha1sum demo.bin download.bin +``` + +```text +d9016032ced6c7515b67a0c556e006c4b25a5858 demo.bin +d9016032ced6c7515b67a0c556e006c4b25a5858 download.bin +``` + +## 5. 验证 RustFS 中的对象 + +列举桶: + +```bash +rc ls rustfs/tus-uploads/ -r +``` + +桶内存放着组装完成的对象,以及每个上传各一个 `.info` 元数据文件——无论进行中还是已完成: + +```text +b7338250daa9a1a79c1343502b57b28f 6 MiB +b7338250daa9a1a79c1343502b57b28f.info 378 B +``` + +对象键即上传 ID,对象内容与上传文件逐字节一致——因此任何 S3 客户端都能直接从桶中读取已完成的上传。 + +![存储在 RustFS 控制台中的 tus 上传](./images/rustfs-tus-uploads.png) + +## 6. 停止或重置 + +保留桶内对象、仅拆除服务器: + +```bash +docker rm -f tusd +``` + +删除已存储的上传: + +```bash +rc rm rustfs/tus-uploads/ --recursive --force +``` + +## 故障排查 + +### `CreateMultipartUpload ... A region must be set when sending requests to S3` + +tusd 从 AWS 环境构建 S3 客户端,端点解析必须提供 region。在凭证之外导出 `AWS_REGION=us-east-1`,如第 1 步所示。 + +### `PATCH` 返回 `404` 或连到了错误的主机 + +`Location` URL 回显的是创建请求的 `Host` 头。当客户端与服务器使用不同主机名(容器名对发布端口)时,应剥掉 URL 中的协议与主机,把路径发给客户端可达的地址。 + +### 服务器重启后上传消失了 + +S3 后端把 `.info` 文件保存在桶里,上传可以跨重启存活。如果有清理任务清空了桶,元数据就会丢失——请保护 `tus-uploads` 前缀不被清理任务触及。 + +## 下一步 + +- 在启用更多 tusd 后端前,先阅读 [S3 兼容性说明](/administration/protocols/s3)。 +- 使用[访问密钥管理](/security-compliance/iam/access-token)创建专用的生产凭证。 +- 按照 [tus 协议文档](https://tus.io/protocols/resumable-upload)了解 tusd 在核心协议之上支持的 creation-with-upload、termination 与 checksum 扩展。 diff --git a/content/zh/developer/integration/reverse-proxy/envoy.md b/content/zh/developer/integration/reverse-proxy/envoy.md new file mode 100644 index 00000000..fc2f2c4c --- /dev/null +++ b/content/zh/developer/integration/reverse-proxy/envoy.md @@ -0,0 +1,192 @@ +--- +title: "Envoy" +description: "在 Envoy 后部署 RustFS,用 TLS 终结的独立路由分别转发 S3 API 与 Console。" +--- + +使用 **Envoy** 终结 TLS,并把独立主机名分别路由到 RustFS 的 S3 API 与 Console。本部署在同一个 Docker 网络上运行 Envoy 和单节点 RustFS。你需要 Docker Engine、两条 DNS 记录,以及一张覆盖两个主机名的 TLS 证书(本指南测试环境使用自签证书)。 + +本指南使用以下示例主机名: + +- `s3.example.com` 对应 S3 API +- `console.example.com` 对应 Console + +请替换为解析到 Docker 主机的主机名。 + +:::warning[Serve S3 from the root path] + +不要把 S3 API 发布在 `/s3/` 之类的路径下。AWS Signature Version 4 的签名包含请求路径与主机,改写其中任何一个都会导致签名失效。Envoy 原样转发传入的 `Host` 头,签名因此保持有效。 + +::: + +## 1. 创建部署目录 + +为 Envoy 配置与 TLS 证书创建目录: + +```bash +mkdir -p rustfs-envoy/certs +cd rustfs-envoy +``` + +本地测试可以生成覆盖两个主机名的自签证书: + +```bash +openssl req -x509 -newkey rsa:2048 -nodes \ + -keyout certs/privkey.pem -out certs/fullchain.pem -days 30 \ + -subj "/CN=*.example.com" \ + -addext "subjectAltName=DNS:s3.example.com,DNS:console.example.com" +chmod 644 certs/privkey.pem +``` + +Envoy 镜像以非 root 用户运行,证书必须可读——`600` 的私钥会产生误导性的 `Failed to load incomplete private key` 报错。 + +## 2. 配置 Envoy + +创建带 HTTPS 监听器和两个虚拟主机的配置。路由超时设为 `timeout: 0s`(禁用),避免长时间的流式 S3 上传被中断: + +```yaml title="envoy.yaml" +static_resources: + listeners: + - name: https + address: {socket_address: {address: 0.0.0.0, port_value: 8443}} + filter_chains: + - transport_socket: + name: envoy.transport_sockets.tls + typed_config: + "@type": type.googleapis.com/envoy.extensions.transport_sockets.tls.v3.DownstreamTlsContext + common_tls_context: + tls_certificates: + - certificate_chain: {filename: /certs/fullchain.pem} + private_key: {filename: /certs/privkey.pem} + filters: + - name: envoy.filters.network.http_connection_manager + typed_config: + "@type": type.googleapis.com/envoy.extensions.filters.network.http_connection_manager.v3.HttpConnectionManager + stat_prefix: rustfs_https + route_config: + virtual_hosts: + - name: s3 + domains: ["s3.example.com", "s3.example.com:*"] + routes: + - match: {prefix: "/"} + route: {cluster: rustfs_s3, timeout: 0s} + - name: console + domains: ["console.example.com", "console.example.com:*"] + routes: + - match: {prefix: "/"} + route: {cluster: rustfs_console, timeout: 0s} + http_filters: + - name: envoy.filters.http.router + typed_config: + "@type": type.googleapis.com/envoy.extensions.filters.http.router.v3.Router + clusters: + - name: rustfs_s3 + connect_timeout: 5s + type: STRICT_DNS + lb_policy: ROUND_ROBIN + load_assignment: + cluster_name: rustfs_s3 + endpoints: + - lb_endpoints: + - endpoint: {address: {socket_address: {address: rustfs, port_value: 9000}}} + - name: rustfs_console + connect_timeout: 5s + type: STRICT_DNS + lb_policy: ROUND_ROBIN + load_assignment: + cluster_name: rustfs_console + endpoints: + - lb_endpoints: + - endpoint: {address: {socket_address: {address: rustfs, port_value: 9001}}} +``` + +`host:*` 形式的 domain 条目很重要:客户端在非标准端口上发送的是 `Host: s3.example.com:8443`,Envoy 按 authority 匹配时包含端口。 + +## 3. 启动 Envoy + +在与 RustFS 相同的 Docker 网络上运行 Envoy,只发布代理端口: + +```bash +docker run -d --name envoy --network oo-rustfs_default -p 8443:8443 \ + -v "$PWD/envoy.yaml":/envoy.yaml:ro \ + -v "$PWD/certs":/certs:ro \ + envoyproxy/envoy:v1.34-latest -c /envoy.yaml +``` + +## 4. 验证两个端点 + +用 `curl --resolve` 把示例主机名指向代理(生产环境由 DNS 完成): + +```bash +curl -sk --resolve s3.example.com:8443:127.0.0.1 \ + https://s3.example.com:8443/health/ready -o /dev/null -w "s3 api: %{http_code}\n" + +curl -sk --resolve console.example.com:8443:127.0.0.1 \ + https://console.example.com:8443/rustfs/console/ -o /dev/null -w "console: %{http_code}\n" +``` + +```text +s3 api: 200 +console: 200 +``` + +`-k` 跳过证书校验,因为这里是自签证书;使用受信任的证书后去掉该参数即可。 + +## 5. 通过 Envoy 发送签名 S3 请求 + +把任意 S3 客户端的端点指向代理即可,如同指向 RustFS。将客户端配置为 `https://s3.example.com:8443` 并启用路径风格寻址;代理证书受信任时,签名后的 AWS Signature Version 4 请求原样透传。要在纯 HTTP 上快速测试,可以增加一个与 HTTPS 监听器共享相同虚拟主机路由的 HTTP 监听器(端口 8080),然后使用以下端点:`http://s3.example.com:8080` + +```bash +rc alias set rustfs-envoy http://s3.example.com:8080 +rc ls rustfs-envoy/rclone-demo/ +``` + +```text +[ ] 0B seed/ +[2026-09-29 11:12:41] 5 B write-test.txt +``` + +由于 Envoy 原样转发 `Host` 头,签名可以在 RustFS 侧通过校验。 + +## 多节点后端 + +分布式 RustFS 部署时,把每个节点加入 S3 cluster: + +```yaml title="envoy.yaml" + - name: rustfs_s3 + connect_timeout: 5s + type: STRICT_DNS + lb_policy: ROUND_ROBIN + load_assignment: + cluster_name: rustfs_s3 + endpoints: + - lb_endpoints: + - endpoint: {address: {socket_address: {address: node1, port_value: 9000}}} + - lb_endpoints: + - endpoint: {address: {socket_address: {address: node2, port_value: 9000}}} + - lb_endpoints: + - endpoint: {address: {socket_address: {address: node3, port_value: 9000}}} + - lb_endpoints: + - endpoint: {address: {socket_address: {address: node4, port_value: 9000}}} +``` + +Console cluster 按相同模式使用 `9001` 端口。 + +## 故障排查 + +### `Failed to load incomplete private key from path` + +Envoy 容器以非 root 用户运行,读不了 `600` 且属主为 root 的私钥。执行 `chmod 644`(或把属主改为容器用户,官方镜像中为 UID `1001`)。 + +### 主机名正确但路由返回 `404` + +Envoy 按 authority(含端口)匹配。如上文配置所示,在每个虚拟主机的 `domains` 里加入 `host:*` 变体。 + +### 经代理访问 RustFS 返回 `Access Denied` + +确认代理没有改写路径或 `Host` 头。签名请求必须以客户端签名时的主机名到达 RustFS。 + +## 下一步 + +- [配置 S3 客户端](/developer/examples/aws-cli) +- [启用虚拟主机风格的桶 URL](/integration/virtual) +- [查看健康与就绪端点](/operations/status-check) diff --git a/content/zh/developer/integration/reverse-proxy/index.md b/content/zh/developer/integration/reverse-proxy/index.md index 3014c0d5..0b0bbdcb 100644 --- a/content/zh/developer/integration/reverse-proxy/index.md +++ b/content/zh/developer/integration/reverse-proxy/index.md @@ -13,6 +13,7 @@ description: "为 RustFS S3 API 和控制台选择并配置反向代理。" - [Traefik](./traefik.md) - [Caddy](./caddy.md) - [HAProxy](./haproxy.md) +- [Envoy](./envoy.md) - [Apache HTTP Server](./httpd.md) ## 相关配置 diff --git a/content/zh/developer/integration/reverse-proxy/meta.json b/content/zh/developer/integration/reverse-proxy/meta.json index f186a2f5..28af26d8 100644 --- a/content/zh/developer/integration/reverse-proxy/meta.json +++ b/content/zh/developer/integration/reverse-proxy/meta.json @@ -4,7 +4,8 @@ "nginx", "traefik", "caddy", + "envoy", "haproxy", "httpd" ] -} \ No newline at end of file +}