diff --git a/README.md b/README.md index 759c51f..d1b33f1 100644 --- a/README.md +++ b/README.md @@ -1,682 +1,212 @@ -# librawssg · [![GitHub tag](https://img.shields.io/github/v/tag/mroczect/librawssg?label=version)](https://github.com/mroczect/librawssg/tags) [![License: MIT](https://img.shields.io/badge/License-MIT-yellow.svg)](https://opensource.org/licenses/MIT) [![CI](https://github.com/mroczect/librawssg/actions/workflows/ci.yml/badge.svg)](https://github.com/mroczect/librawssg/actions/workflows/ci.yml) - -**librawssg** is the engine‑agnostic, safety‑first kernel for building static site generators in Rust. -It gives you all the primitives you need: filesystem abstraction, frontmatter parsing, Markdown rendering, template rendering, content processing pipelines, feed & sitemap generation, and a secure development server. - -The library does **not** include a CLI – you write your own `main.rs` and compose the parts you need. -Optional built‑in implementations for **Tera** and **pulldown‑cmark** are available behind feature flags. - ---- - -## Table of Contents - -- [What's New in v0.5.0](#whats-new-in-v050) -- [Installation](#installation) -- [Quick Start](#quick-start) -- [Architecture](#architecture) -- [API Reference](#api-reference) - - [Configuration](#configuration) - - [YAML Configuration File](#yaml-configuration-file) - - [ConfigLoader trait](#configloader-trait) - - [RawssgConfig::validate](#rawssgconfigvalidate) - - [RawssgConfig default](#rawssgconfig-default) - - [Error Handling](#error-handling) - - [Filesystem Abstraction](#filesystem-abstraction) - - [FileSystem trait](#filesystem-trait) - - [RealFs](#realfs) - - [Implementing a Custom FileSystem](#implementing-a-custom-filesystem) - - [Markdown Rendering](#markdown-rendering) - - [MarkdownRenderer trait](#markdownrenderer-trait) - - [PulldownMarkdown](#pulldownmarkdown) - - [Implementing a Custom Markdown Renderer](#implementing-a-custom-markdown-renderer) - - [Template Rendering](#template-rendering) - - [TemplateRenderer trait](#templaterenderer-trait) - - [Context trait](#context-trait) - - [TeraRenderer](#terarenderer) - - [Implementing a Custom Template Engine](#implementing-a-custom-template-engine) - - [Content Pipeline](#content-pipeline) - - [ContentHandler trait](#contenthandler-trait) - - [MarkdownPageHandler](#markdownpagehandler) - - [StaticFileHandler](#staticfilehandler) - - [build_page_context](#build_page_context) - - [Adding Custom Handlers](#adding-custom-handlers) - - [Site Builder](#site-builder) - - [SiteBuilder](#sitebuilder) - - [Site](#site) - - [Atomic Generation & Cross-Device Fallback](#atomic-generation--cross-device-fallback) - - [Feed & Sitemap](#feed--sitemap) - - [generate_feed and generate_sitemap](#generate_feed-and-generate_sitemap) - - [Context Builders](#context-builders) - - [Utility Functions](#utility-functions) - - [safe_path](#safe_path) - - [slugify](#slugify) - - [relative_prefix](#relative_prefix) - - [match_pattern](#match_pattern) - - [Type Reference](#type-reference) - - [RawssgConfig](#rawssgconfig) - - [GlobalConfig](#globalconfig) - - [BuildConfig](#buildconfig) - - [ContentTypeDef](#contenttypedef) - - [GeneratorsConfig & GeneratorDef](#generatorsconfig--generatordef) - - [NavItem](#navitem) - - [PageFrontMatter](#pagefrontmatter) - - [PageContext](#pagecontext) - - [Dev Server & Watcher (serve feature)](#dev-server--watcher-serve-feature) -- [Feature Flags](#feature-flags) -- [Security](#security) -- [Full Customisation](#full-customisation) - - [Step‑by‑Step: Building a Fully Custom SSG](#stepbystep-building-a-fully-custom-ssg) -- [Testing](#testing) -- [Contributing](#contributing) -- [License](#license) - ---- - -## What's New in v0.5.0 - -- **Robust path security** – Symlink‑safe output writing via canonicalised parent directories. -- **Correct glob matching** – `**` patterns now match exactly as expected (e.g., `blog/**/*.html` no longer matches `.md`). -- **Completely trait‑based** – Added `rename` to `FileSystem`; all I/O goes through the trait for full mockability. -- **Better defaults** – A default `page` content type (`**/*.md` → `base.html`) is included out‑of‑the‑box. -- **Optional context builders** – Feed and sitemap context builders are now only required when the respective generator is enabled. -- **Clearer error messages** – Missing closing `---` in frontmatter is reported explicitly; config loading failures are logged. -- **Improved testing** – Property‑based tests, dynamic server ports, and a complete mock filesystem. - ---- - -## Installation - -### Method 1: `cargo add` (Git dependency – recommended) +# librawssg -```bash -cargo add --git https://github.com/mroczect/librawssg.git --tag v0.5.0 librawssg -cargo add --git https://github.com/mroczect/librawssg.git --tag v0.5.0 librawssg --features tera,pulldown -``` +A modular static site generator library for Rust. -### Method 2: Manual `Cargo.toml` entry +`librawssg` is a collection of crates that together form a flexible and extensible framework for building static site generators. The project is designed with modularity, testability, and safety in mind, leveraging Rust's type system and trait abstractions. -```toml -[dependencies] -librawssg = { git = "https://github.com/mroczect/librawssg.git", tag = "v0.5.0" } -librawssg = { git = "https://github.com/mroczect/librawssg.git", tag = "v0.5.0", features = ["tera", "pulldown"] } -``` +## Features -### Method 3: Path dependency (local development) +- **Modular architecture** – Each aspect (configuration, filesystem, content processing, templating, compilation) is isolated into its own crate. +- **Pluggable processors** – Define custom content processors via the `Processor` trait. +- **Template engine integration** – Built-in support for Tera templates through the `TeraRenderer` (optional, enabled by default). +- **Strong filesystem abstraction** – Trait-based filesystem with built-in path traversal protection. +- **Atomic output generation** – The build pipeline writes to a temporary directory and atomically replaces the final output. +- **Comprehensive configuration** – YAML/JSON support, validation, and nested site/build settings. +- **Extensible** – Add custom renderers, context builders, and post-processing generators. +- **Strict linting** – Deny-level lints for clippy and rustc ensure high code quality. +- **Demo application** – A complete example showing how to assemble the parts into a working static site. -```bash -git clone https://github.com/mroczect/librawssg.git -cd librawssg -# in your project's Cargo.toml: -librawssg = { path = "../librawssg", features = ["tera", "pulldown"] } -``` +## Repository Structure ---- +The workspace consists of the following crates: -## Quick Start +| Crate | Description | +| --------------------- | -------------------------------------------------------------- | +| `librawssg` | Facade crate that re-exports all other crates for convenience. | +| `librawssg_config` | Configuration data structures and validation. | +| `librawssg_fs` | Filesystem abstraction trait and real implementation. | +| `librawssg_handler` | Core document, metadata, and processor contracts. | +| `librawssg_templates` | Rendering traits and Tera implementation. | +| `librawssg_compiler` | Build pipeline orchestration. | +| `librawssg_error` | Unified error types and result alias. | +| `librawssg_demo` | Example application demonstrating usage of the framework. | -```rust -use librawssg::SiteBuilder; -use librawssg::site::TeraRenderer; -use librawssg::markdown::PulldownMarkdown; +## Getting Started -fn main() -> Result<(), Box> { - let mut tera = TeraRenderer::new(); - tera.add_raw_template("base.html", "{{ page_content }}")?; - let md = PulldownMarkdown; +### Prerequisites - let mut config = librawssg::RawssgConfig::default(); - config.build.content_dir = "content".into(); - config.build.output_dir = "dist".into(); - - let site = SiteBuilder::new() - .config(config) - .with_template_renderer(Box::new(tera)) - .with_markdown_renderer(Box::new(md)) - .build()?; - - site.generate()?; - Ok(()) -} -``` - ---- - -## Architecture - -``` -src/ - config/ ConfigLoader trait, YamlConfigLoader, DefaultConfig - error.rs RawssgError (miette + thiserror) - frontmatter.rs YAML frontmatter extraction and Markdown rendering - fs/ FileSystem trait, RealFs - markdown.rs MarkdownRenderer trait, optional PulldownMarkdown - serve/ Dev server and file watcher (feature "serve") - site/ - builders/ Site and SiteBuilder structs - context.rs FeedContextBuilder, SitemapContextBuilder traits - feed.rs generate_feed - mod.rs Core traits: TemplateRenderer, Context, ContentHandler - page.rs build_page_context - sitemap.rs generate_sitemap - types.rs All configuration and page context types - util.rs safe_path, slugify, relative_prefix, match_pattern -``` +- Rust toolchain (stable, edition 2024) – install via [rustup](https://rustup.rs/) +- Cargo (comes with Rust) ---- +### Building the Project -## API Reference +Clone the repository and build all crates: -### Configuration - -#### YAML Configuration File - -```yaml -site: - site_name: "My Site" - description: "A blog about Rust" - base_url: "https://example.com" - language: "en" - author: "Alice" -build: - content_dir: content - output_dir: dist - templates_dir: templates - static_dir: static -content_types: - - name: blog - pattern: blog/**/*.md - template: post.html - list_template: blog_list.html - list_enabled: true - - name: page - pattern: **/*.md - template: page.html -generators: - rss: - enabled: true - path: feed.xml - template: rss.xml - sitemap: - enabled: true - path: sitemap.xml - template: sitemap.xml -``` - -#### ConfigLoader trait - -```rust -pub trait ConfigLoader: Send + Sync { - fn load(&self) -> Result; - fn load_or_default(&self) -> RawssgConfig; -} -``` - -- `YamlConfigLoader>` – reads a YAML file. -- `DefaultConfig` – returns `RawssgConfig::default()`. - -#### RawssgConfig::validate - -- Checks that `site_name` is non‑empty. -- At least one `content_types` entry must exist. -- Each content type must have a valid glob pattern and a template name. -- If RSS or sitemap is enabled, their `path` and `template` must be set. - -#### RawssgConfig default - -Since v0.5.0, the default configuration includes a single content type: - -```rust -ContentTypeDef { - name: "page".into(), - pattern: "**/*.md".into(), - template: "base.html".into(), - list_template: None, - list_enabled: false, -} -``` - -You can remove it with `config.content_types.clear()` and define your own. - -### Error Handling - -`RawssgError` implements `std::error::Error`, `Display`, and `miette::Diagnostic`. - -```rust -match err { - RawssgError::Frontmatter { path, source } => { /* ... */ } - RawssgError::PathTraversal(msg) => { /* ... */ } - // ... -} -``` - -### Filesystem Abstraction - -#### FileSystem trait - -```rust -pub trait FileSystem: Send + Sync { - fn read_to_string(&self, path: &Path) -> io::Result; - fn read_bytes(&self, path: &Path) -> io::Result>; - fn write(&self, path: &Path, content: &[u8]) -> io::Result<()>; - fn create_dir_all(&self, path: &Path) -> io::Result<()>; - fn remove_dir_all(&self, path: &Path) -> io::Result<()>; - fn exists(&self, path: &Path) -> bool; - fn is_dir(&self, path: &Path) -> bool; - fn is_file(&self, path: &Path) -> bool; - fn read_dir(&self, path: &Path) -> io::Result>; - fn copy_file(&self, from: &Path, to: &Path) -> io::Result; - fn walk_dir(&self, root: &Path) -> io::Result>; - fn canonicalize(&self, path: &Path) -> io::Result; - fn rename(&self, from: &Path, to: &Path) -> io::Result<()>; // new in v0.5.0 -} -``` - -#### RealFs - -Default implementation delegating to `std::fs` and `walkdir`. - -#### Implementing a Custom FileSystem - -```rust -struct MyFs; -impl FileSystem for MyFs { - // implement all methods; e.g., read from database or network - fn read_to_string(&self, path: &Path) -> io::Result { /* ... */ } - // ... etc. -} -let site = SiteBuilder::new().with_fs(Box::new(MyFs)).build()?; +```bash +git clone https://github.com/mroczect/librawssg.git +cd librawssg +cargo build ``` -### Markdown Rendering +### Running the Demo -#### MarkdownRenderer trait +The `librawssg_demo` crate provides a working example. To run it: -```rust -pub trait MarkdownRenderer: Send + Sync { - fn render(&self, markdown: &str) -> String; -} +```bash +cargo run -p librawssg_demo ``` -#### PulldownMarkdown - -Available with `pulldown` feature. Enables tables, strikethrough, task lists. - -#### Implementing a Custom Markdown Renderer - -```rust -struct MyMd; -impl MarkdownRenderer for MyMd { - fn render(&self, md: &str) -> String { my_parser(md) } -} -``` +This will process `.raw` HTML fragment files from `librawssg_demo/src/content`, render them using a Tera template, copy static assets, and output the site into `librawssg_demo/dist`. -### Template Rendering +### Using as a Library -#### TemplateRenderer trait +Add `librawssg` to your `Cargo.toml`: -```rust -pub trait TemplateRenderer: Send + Sync { - fn render(&self, template_name: &str, context: &dyn Context) -> Result; -} +```toml +[dependencies] +librawssg = "1.0.0" ``` -#### Context trait +Then you can import the necessary components. Here is a minimal example that sets up a pipeline: ```rust -pub trait Context: Send + Sync { - fn as_any(&self) -> &dyn Any; - fn as_mut_any(&mut self) -> &mut dyn Any; -} -``` +use librawssg::{ + Config, ContentRule, Document, FileSystem, Metadata, PipelineBuilder, Processor, + RealFs, RenderContext, Renderer, TeraContextBuilder, TeraRenderer, +}; +use std::path::{Path, PathBuf}; -#### TeraRenderer - -- `new()` – creates an empty Tera instance. -- `add_raw_template(name, content)` – registers an inline template. - -#### Implementing a Custom Template Engine - -Implement `TemplateRenderer` and a `Context` wrapper. -Example: MiniJinja. - -```rust -struct MiniJinjaRenderer { env: mini_jinja::Environment<'static> } -impl TemplateRenderer for MiniJinjaRenderer { - fn render(&self, name: &str, ctx: &dyn Context) -> Result { - let tmpl = self.env.get_template(name).map_err(|e| RawssgError::Template(e.to_string()))?; - let data = ctx.as_any().downcast_ref::().unwrap(); - tmpl.render(data).map_err(|e| RawssgError::Template(e.to_string())) +// Implement a custom processor for .txt files +struct TextProcessor; +impl Processor for TextProcessor { + fn name(&self) -> &'static str { "text" } + fn can_process(&self, rel: &Path, _orig: &Path) -> bool { + rel.extension().and_then(|e| e.to_str()) == Some("txt") } -} -impl Context for serde_json::Value { /* as_any downcast */ } -``` - -### Content Pipeline - -#### ContentHandler trait - -```rust -pub trait ContentHandler: Send + Sync { - fn can_handle(&self, relative_path: &Path, original_path: &Path) -> bool; - fn process(&self, fs: &dyn FileSystem, md_renderer: &dyn MarkdownRenderer, - file_path: &Path, content_dir: &Path) -> Result, RawssgError>; -} -``` - -Return `None` to skip a file. - -#### MarkdownPageHandler - -Handles `.md` files; extracts frontmatter, renders Markdown. - -#### StaticFileHandler - -Always returns `None` (catch‑all, non‑Markdown files become static assets). - -#### build_page_context - -```rust -pub fn build_page_context(fs: &dyn FileSystem, md_renderer: &dyn MarkdownRenderer, - file_path: &Path, content_dir: &Path) -> Result, RawssgError>; -``` - -Skips drafts. Returns a `PageContext` with URL, depth, date formatting. - -#### Adding Custom Handlers - -```rust -struct AsciiDocHandler; -impl ContentHandler for AsciiDocHandler { - fn can_handle(&self, _rel: &Path, orig: &Path) -> bool { - orig.extension().map_or(false, |e| e == "adoc") - } - fn process(&self, fs: &dyn FileSystem, _md: &dyn MarkdownRenderer, ...) -> Result, RawssgError> { - let content = fs.read_to_string(file_path)?; - let html = asciidoc_render(&content); - Ok(Some(PageContext { content_html: html, .. })) + fn process( + &self, + fs: &dyn FileSystem, + rel: &Path, + content_dir: &Path, + ) -> librawssg::Result> { + let body = fs.read_to_string(&content_dir.join(rel))?; + let meta = Metadata::new("Page", "Description")?; + let url = rel.with_extension("html").to_string_lossy().to_string(); + let doc = Document::new( + meta, + body, + url.clone(), + PathBuf::from(&url), + rel.to_path_buf(), + 0, + "page".to_string(), + false, + )?; + Ok(Some(doc)) } } -let builder = SiteBuilder::new().add_handler(Box::new(AsciiDocHandler)); -``` - -### Site Builder - -#### SiteBuilder - -```rust -SiteBuilder::new() - .config(config) - .load_config("config.yml")? // alternative to .config() - .content_dir("my_content") - .output_dir("public") - .with_fs(Box::new(RealFs)) - .with_markdown_renderer(Box::new(PulldownMarkdown)) - .with_template_renderer(Box::new(TeraRenderer::new())) - .with_feed_context_builder(Box::new(TeraFeedContextBuilder)) - .with_sitemap_context_builder(Box::new(TeraSitemapContextBuilder)) - .add_handler(Box::new(MyHandler)) - .build()?; -``` - -- `content_dir`, `output_dir` can be overridden by config values if left as default (`"content"`, `"dist"`). -- `feed_context_builder` and `sitemap_context_builder` are only required when the corresponding generator is enabled. - -#### Site - -```rust -let pages: &[PageContext] = site.pages(); -site.generate()?; // atomic write to output_dir -``` - -`generate()`: - -1. Writes all pages (HTML) to `output_dir`. -2. Copies static assets from `static_dir`. -3. Copies non‑Markdown files from `content_dir`. -4. Optionally generates RSS and sitemap (if `tera` feature + enabled). -5. Uses atomic write: temp dir → rename (with cross‑device fallback). - -#### Atomic Generation & Cross-Device Fallback - -If `rename` fails with `CrossesDevices`, the library performs a recursive copy and then deletes the temporary directory. -### Feed & Sitemap - -#### generate_feed and generate_sitemap - -```rust -pub fn generate_feed(renderer: &dyn TemplateRenderer, config: &RawssgConfig, - posts: &[&PageContext], base_url: &str, context_builder: &dyn FeedContextBuilder) -> Result; -pub fn generate_sitemap(renderer: &dyn TemplateRenderer, config: &RawssgConfig, - pages: &[PageContext], base_url: &str, context_builder: &dyn SitemapContextBuilder) -> Result; -``` +fn main() -> Result<(), Box> { + let mut config = Config::new().with_site_name("My Site"); + config.add_content_rule(ContentRule::new("page", "**/*.txt", "base.tera")); + config.build.content_dir = "content".into(); + config.build.output_dir = "dist".into(); + config.build.static_dir = "static".into(); -Callers are responsible for writing the returned string to the output file; `Site::generate` does this automatically. + let mut renderer = TeraRenderer::new(); + renderer.load_templates_dir(Path::new("templates"))?; -#### Context Builders + let pipeline = PipelineBuilder::new() + .config(config) + .content_dir("content") + .output_dir("dist") + .with_fs(Box::new(RealFs)) + .with_renderer(Box::new(renderer)) + .with_context_builder(Box::new(TeraContextBuilder)) + .add_processor(Box::new(TextProcessor)) + .build()?; -```rust -pub trait FeedContextBuilder: Send + Sync { - fn build_feed_context(&self, config: &RawssgConfig, posts: &[&PageContext], base_url: &str) - -> Result, RawssgError>; + pipeline.run()?; + println!("Site generated!"); + Ok(()) } ``` -Default Tera implementations insert `site`, `posts`/`pages`, and `base_url`. Custom builders can add extra variables (e.g., `ctx.insert("custom", &"value")`). - -### Utility Functions - -#### safe_path - -```rust -pub fn safe_path(fs: &dyn FileSystem, base: &Path, candidate: &Path) -> Result; -``` - -- Canonicalises `base`. -- For existing files: canonicalises the candidate, checks it stays inside `base`. -- For non‑existent files (output): canonicalises the parent directory and verifies confinement. -- Returns an absolute, safe path. - -#### slugify - -Converts a string to lowercase, alphanumeric + hyphens. Example: `"Hello World!"` → `"hello-world"`. - -#### relative_prefix - -Returns `"./"` for depth 0, `"../"` repeated for deeper paths. - -#### match_pattern - -Glob matching with `*` (single segment) and `**` (multi‑segment). Correctly handles patterns like `blog/**/*.html`. - -### Type Reference +For a more detailed example, see the `librawssg_demo` source code. -#### RawssgConfig +## Core Concepts -```rust -pub struct RawssgConfig { - pub site: GlobalConfig, - pub build: BuildConfig, - pub content_types: Vec, - pub generators: GeneratorsConfig, -} -``` +### Configuration -- `validate()` ensures invariants. -- `Default` now includes one content type. +The `Config` struct holds all settings required for the build. It includes: -#### GlobalConfig +- `site`: Site-wide metadata (name, description, navigation, etc.) +- `build`: Paths for content, output, templates, and static assets. +- `content_rules`: A list of `ContentRule` objects that map file patterns to templates. +- `extra`: Arbitrary key-value data. -```rust -pub struct GlobalConfig { - pub site_name: String, // default "rawssg" - pub description: Option, - pub language: Option, // default Some("en") - pub base_url: Option, - pub author: Option, - pub repo_url: Option, - pub license: Option, - pub navbar: Vec, - pub sidebar: Vec, -} -``` +Configuration can be loaded from YAML or JSON using `Config::from_yaml_str` / `Config::from_json_str`. -#### BuildConfig +### Content Processing -```rust -pub struct BuildConfig { - pub content_dir: String, // default "content" - pub output_dir: String, // default "dist" - pub templates_dir: String, // default "templates" - pub static_dir: String, // default "static" -} -``` +Content files are processed by implementations of the `Processor` trait. Each processor declares which files it can handle via `can_process()`, and then transforms them into `Document` objects. The pipeline walks the content directory, determines the appropriate processor for each file, and collects the resulting documents. -#### ContentTypeDef +### Rendering -```rust -pub struct ContentTypeDef { - pub name: String, - pub pattern: String, // glob - pub template: String, - pub list_template: Option, - pub list_enabled: bool, -} -``` +The `Renderer` trait abstracts template rendering. The built-in `TeraRenderer` uses the Tera template engine. A `ContextBuilder` creates the render context for each document; the default `TeraContextBuilder` populates it with page and site data. -#### GeneratorsConfig & GeneratorDef +### Build Pipeline -```rust -pub struct GeneratorsConfig { pub rss: GeneratorDef, pub sitemap: GeneratorDef } -pub struct GeneratorDef { - pub enabled: bool, // default false - pub path: String, - pub template: String, -} -``` +The `PipelineBuilder` assembles all components (filesystem, renderer, processors, context builder, generators) and produces a `Pipeline`. Calling `pipeline.run()` performs the following steps: -#### NavItem +1. Processes all content files. +2. Renders non-list documents. +3. Optionally generates list pages (index pages) for content types with list support enabled. +4. Copies static assets. +5. Executes any custom generators. +6. Atomically replaces the output directory. -```rust -pub struct NavItem { pub label: String, pub url: String } -``` +## Customization -#### PageFrontMatter +You can extend the framework by implementing the following traits: -```rust -pub struct PageFrontMatter { - pub title: String, - pub desc: String, - pub author: Option, - pub date: Option, - pub tags: Vec, - pub draft: bool, - // ... -} -``` +- **`Processor`** – For handling new file types or custom transformations. +- **`Renderer`** and **`RenderContext`** – To integrate a different template engine. +- **`ContextBuilder`** – To customize the data passed to templates. +- **`Generator`** – To add extra outputs like RSS feeds, sitemaps, or search indexes. -#### PageContext +All components are passed to the pipeline as boxed trait objects, so they are easily swappable. -```rust -pub struct PageContext { - pub frontmatter: PageFrontMatter, - pub content_html: String, - pub url: String, - pub file_path: String, - pub depth: usize, - pub pub_date: Option, - pub content_type: String, - pub is_list: bool, - pub list_items: Option>, -} -``` +## Development -### Dev Server & Watcher (serve feature) +### Workspace Lints -```rust -use librawssg::serve::start_dev_server; -start_dev_server(Path::new("dist"), 8080)?; -``` +The workspace enforces strict linting via `[workspace.lints]` in the root `Cargo.toml`. Many clippy and rustc lints are set to `deny`, including `unsafe_code = "forbid"`, `unwrap_used = "deny"`, `expect_used = "deny"`, `panic = "deny"`, and many others. This ensures high code quality and safety. When contributing, please ensure your code passes `cargo clippy --all --all-targets --all-features -- -D warnings` and `cargo fmt --check`. -- Serves files with correct MIME types (including `text/plain` for `.txt`, `.md`, `.yaml`). -- 404 for missing, 500 for internal errors. +### Running Tests -```rust -use librawssg::serve::watch_dirs; -let _watcher = watch_dirs(&[content_path, templates_path], || { rebuild(); })?; +```bash +cargo test --workspace ``` -Uses `notify` to trigger on `Modify`, `Create`, `Remove`. - ---- - -## Feature Flags - -| Feature | Deps | Description | -| ---------- | -------------------- | -------------------------------- | -| `tera` | `tera` | `TeraRenderer`, context builders | -| `pulldown` | `pulldown-cmark` | `PulldownMarkdown` | -| `serve` | `tiny_http`,`notify` | Dev server + file watcher | - -All disabled by default. - ---- - -## Security - -- **Path confinement**: `safe_path` prevents directory traversal for both existing and new files. -- **Atomic output**: Temp directory → rename; fallback copy‑and‑delete ensures atomicity across devices. -- **Configuration validation**: All YAML keys are known; glob patterns are validated. -- **Trait‑based I/O**: Every disk access goes through `FileSystem`, allowing sandboxing and auditing. - ---- - -## Full Customisation - -Every core component is a trait. You can: - -- **Filesystem**: Implement `FileSystem` to read from database, in‑memory store, or network. -- **Markdown**: Any parser via `MarkdownRenderer`. -- **Templates**: Any engine via `TemplateRenderer` + `Context`. -- **Content handlers**: Add new file processors (e.g., AsciiDoc, reStructuredText). -- **Feed/Sitemap**: Custom context builders inject arbitrary variables. -- **Site generation**: Use `SiteBuilder` to build `Site`, then replace `generate()` with your own logic. - -### Step‑by‑Step: Building a Fully Custom SSG - -1. **Define your custom types** – implement the required traits. -2. **Load configuration** – use `YamlConfigLoader` or build `RawssgConfig` programmatically. -3. **Instantiate `SiteBuilder`** with your custom implementations. -4. **Call `.build()`** to obtain a `Site`. -5. **Generate** using `site.generate()`, or iterate over `site.pages()` for custom output. - ---- - -## Testing +### Formatting ```bash -cargo fmt --all -- --check -cargo clippy --all-targets --all-features -- -D warnings -cargo test --workspace --all-features +cargo fmt --all ``` -The test suite includes: - -- `MockFs` – full mock filesystem (with `rename`). -- Mock renderers. -- Property‑based tests (`proptest`). -- Integration tests: full generation, drafts, blog lists, RSS/sitemap errors. -- Dynamic port allocation for server tests. +## Contributing ---- +Contributions are welcome! Please read `CONTRIBUTING.md` and `CODE_OF_CONDUCT.md` for guidelines. By participating, you agree to abide by the project's code of conduct. -## Contributing +## License -Please read [`CONTRIBUTING.md`](CONTRIBUTING.md) for guidelines. -All contributions are welcome – issues, PRs, documentation improvements. +This project is licensed under the **MIT License**. See the [LICENSE](LICENSE) file for details. ---- +## Acknowledgements -## License +This project uses the following open-source crates (among others): -MIT. See [`LICENSE`](LICENSE). +- [Tera](https://github.com/Keats/tera) – Template engine +- [Serde](https://serde.rs/) – Serialization framework +- [WalkDir](https://github.com/BurntSushi/walkdir) – Directory traversal diff --git a/librawssg/README.md b/librawssg/README.md index 6efdc79..9388444 100644 --- a/librawssg/README.md +++ b/librawssg/README.md @@ -2,11 +2,442 @@ A modular static site generator library for Rust. -This facade crate re-exports the essential building blocks: -- Configuration types -- Filesystem abstraction -- Content processor contracts -- Template rendering (Tera built-in) -- Build pipeline orchestration - -See the crate documentation for usage examples. +This facade crate re‑exports the essential building blocks from the `librawssg` ecosystem, providing a single convenient entry point for building static site generators. It aggregates configuration management, filesystem abstraction, content processing, template rendering (with Tera built‑in), and build pipeline orchestration. + +--- + +## Table of Contents + +1. [Overview](#overview) +2. [Installation](#installation) +3. [Modules](#modules) +4. [Core Types and Traits](#core-types-and-traits) + - [Configuration](#configuration) + - [Filesystem](#filesystem) + - [Content Handling](#content-handling) + - [Template Rendering](#template-rendering) + - [Compiler Pipeline](#compiler-pipeline) + - [Error Handling](#error-handling) +5. [Usage Example](#usage-example) +6. [Full API Reference](#full-api-reference) + - [Configuration Types](#configuration-types) + - [Filesystem Types](#filesystem-types) + - [Handler Types](#handler-types) + - [Template Types](#template-types) + - [Compiler Types](#compiler-types) + - [Error Types](#error-types) +7. [Feature Flags](#feature-flags) +8. [License](#license) + +--- + +## Overview + +`librawssg` is the top‑level crate that brings together six specialized crates: + +- **`librawssg_config`** – Configuration data structures and validation. +- **`librawssg_fs`** – Trait‑based filesystem abstraction with path traversal protection. +- **`librawssg_handler`** – Core document and metadata types, plus the `Processor` trait. +- **`librawssg_templates`** – Rendering traits (`Renderer`, `RenderContext`) and a Tera implementation. +- **`librawssg_compiler`** – Build pipeline orchestration. +- **`librawssg_error`** – Unified error enum and `Result` alias. + +By depending on `librawssg`, you get all these components without needing to specify each one individually. The facade also re‑exports the most commonly used types at the crate root for ergonomic access. + +--- + +## Installation + +Add `librawssg` to your `Cargo.toml`: + +```toml +[dependencies] +librawssg = "1.0.0" +``` + +If you are working in the same workspace as the `librawssg` source, you can use a path dependency: + +```toml +[dependencies] +librawssg = { path = "../librawssg" } +``` + +The crate is compatible with Rust edition 2024 and later. + +--- + +## Modules + +The crate organises its re‑exports into submodules for clarity: + +- **`librawssg::config`** – Configuration types (`Config`, `SiteConfig`, `BuildConfig`, `ContentRule`, `NavItem`). +- **`librawssg::fs`** – Filesystem trait and `RealFs`. +- **`librawssg::handler`** – Document, metadata, and processor contracts. +- **`librawssg::templates`** – Rendering traits and `TeraRenderer`. +- **`librawssg::compiler`** – Pipeline builder, pipeline, context builders, generators. +- **`librawssg::error`** – Error type and `Result` alias. + +Additionally, the most important types are also re‑exported directly at the crate root for convenience. + +--- + +## Core Types and Traits + +### Configuration + +The configuration system revolves around the `Config` struct, which contains site settings, build paths, and content processing rules. + +- **`Config`** – Top‑level configuration. + - **Fields**: `site: SiteConfig`, `build: BuildConfig`, `content_rules: Vec`, `extra: HashMap`. + - **Constructors**: + - `Config::new() -> Config` + - `Config::default() -> Config` + - **Builder method**: `with_site_name(name: impl Into) -> Self` + - **Rule management**: + - `add_content_rule(&mut self, rule: ContentRule)` + - `find_rule_by_name(&self, name: &str) -> Option<&ContentRule>` + - `remove_rule_by_name(&mut self, name: &str) -> Option` + - `has_duplicate_rule_names(&self) -> bool` + - **Validation**: `validate(&self) -> Result<()>` + - **Serialization**: + - `from_yaml_str(yaml: &str) -> Result` + - `to_yaml_string(&self) -> Result` + - `from_json_str(json: &str) -> Result` + - `to_json_string(&self) -> Result` + +- **`SiteConfig`** – Global site metadata. + - **Fields**: `navbar`, `sidebar`, `site_name`, `description`, `language`, `base_url`, `author`, `repo_url`, `license`, `extra`. + - **Constructors**: `SiteConfig::new(site_name: impl Into) -> Self`, `SiteConfig::default()`. + - **Defaults**: `site_name = "librawssg"`, `language = Some("en")`. + +- **`BuildConfig`** – Filesystem path settings. + - **Fields**: `content_dir`, `output_dir`, `templates_dir`, `static_dir`. + - **Constructors**: `BuildConfig::new()`, `BuildConfig::default()`. + - **Defaults**: `"content"`, `"dist"`, `"templates"`, `"static"`. + +- **`ContentRule`** – Defines how a group of files should be processed. + - **Fields**: `name`, `pattern`, `template`, `list_template: Option`, `list_enabled: bool`, `extra: HashMap`. + - **Constructor**: `ContentRule::new(name, pattern, template) -> Self`. + - **Default**: all strings empty, `list_enabled = false`. + +- **`NavItem`** – Navigation menu entry. + - **Fields**: `label: String`, `url: String`, `children: Vec`. + - **Constructor**: `NavItem::new(label, url) -> Self`. + +### Filesystem + +- **`FileSystem` trait** – Abstract filesystem operations. All methods return `io::Result` or `bool`. + - Required methods: + - `read_to_string`, `read_bytes`, `write`, `create_dir_all`, `remove_dir_all`, `remove_file`, `create_dir`, `exists`, `is_dir`, `is_file`, `read_dir`, `copy_file`, `copy_dir_all`, `walk_dir`, `canonicalize`, `rename`, `atomic_write`, `touch`, `metadata`, `symlink_metadata`, `permissions`, `set_permissions`, `read_link`, `hard_link`. + - Provided methods (with default implementations): + - `is_symlink` + - `canonicalize_or_join` + - `safe_join` – **Important for security**: ensures the resulting path stays within a base directory. + - `copy` + - `rename_or_copy` + +- **`RealFs`** – Zero‑sized struct implementing `FileSystem` using `std::fs` and `walkdir`. + +### Content Handling + +- **`Document`** – Represents a processed content item. + - **Fields**: `metadata: Metadata`, `body: String`, `url: String`, `output_path: PathBuf`, `source_path: PathBuf`, `depth: usize`, `content_type: String`, `is_list: bool`, `list_items: Option>`, `taxonomies: HashMap>`. + - **Constructor**: `Document::new(metadata, body, url, output_path, source_path, depth, content_type, is_list) -> Result`. + - **Methods**: + - `relative_url(&self) -> &str` + - `add_taxonomy(&mut self, name: impl Into, items: Vec)` + - `depth(&self) -> usize` + - `with_list_items(self, items: Vec) -> Self` + +- **`Metadata`** – Front matter data. + - **Fields**: `title`, `description`, `author: Option`, `repo_url`, `license`, `date: Option`, `updated: Option`, `tags: Vec`, `draft: bool`, `extra: HashMap`. + - **Constructor**: `Metadata::new(title, description) -> Result`. + - **Methods**: + - `is_draft(&self) -> bool` + - `insert_extra(&mut self, key, value)` + - `get_extra(&self, key: &str) -> Option<&serde_json::Value>` + +- **`Processor` trait** – Interface for transforming source files into `Document`s. + - Required methods: + - `name(&self) -> &str` + - `can_process(&self, relative_path: &Path, original_path: &Path) -> bool` + - `process(&self, fs: &dyn FileSystem, relative_path: &Path, content_dir: &Path) -> Result>` + - Provided method: `priority(&self) -> i32` (default 0). + +### Template Rendering + +- **`RenderContext` trait** – Type‑erased context for renderers. + - Required methods: `as_any(&self) -> &dyn Any`, `as_mut_any(&mut self) -> &mut dyn Any`. + +- **`Renderer` trait** – Interface for template rendering. + - Required method: `render(&self, template_name: &str, context: &dyn RenderContext) -> Result`. + +- **`TeraRenderer`** – Concrete renderer using the Tera template engine. + - Available when the `tera` feature is enabled (enabled by default). + - **Constructors**: `TeraRenderer::new()`, `Default`. + - **Methods**: + - `add_raw_template(&mut self, name: &str, content: &str) -> Result<()>` + - `add_template_file(&mut self, path: &Path) -> Result<()>` + - `add_template_files_from_dir(&mut self, dir: &Path) -> Result<()>` + - `load_templates_dir(&mut self, dir: &Path) -> Result<()>` + - `enable_autoescape(&mut self)` + - `render_str(&self, template_str: &str, context: &dyn RenderContext) -> Result` + - `as_tera(&self) -> &tera::Tera` + - `as_tera_mut(&mut self) -> &mut tera::Tera` + +### Compiler Pipeline + +- **`PipelineBuilder`** – Builds a `Pipeline`. + - **Constructor**: `PipelineBuilder::new()`. + - **Builder methods**: `config`, `load_config`, `content_dir`, `output_dir`, `with_fs`, `with_renderer`, `add_processor`, `with_context_builder`, `add_generator`. + - **Build**: `build(self) -> Result`. + +- **`Pipeline`** – Executes the site generation. + - **Method**: `run(&self) -> Result<()>`. + - **Accessor**: `config(&self) -> &Config`. + +- **`ContextBuilder` trait** – Creates a `RenderContext` from `Config` and `Document`. + - Required method: `build_context(&self, config: &Config, doc: &Document) -> Result>`. + +- **`TeraContextBuilder`** – Default implementation that produces a `tera::Context` with common page and site variables. + +- **`Generator` trait** – Custom post‑processing step. + - Required method: `generate(&self, pipeline: &Pipeline, output_base: &Path) -> Result<()>`. + +- **`match_pattern` function** (in `compiler::pattern`) – Glob matching for content rule patterns. + +### Error Handling + +- **`Error` enum** – Variants: `Io`, `Config`, `Metadata`, `Render`, `Processor`, `Generator`, `PathTraversal`, `MissingConfig`, `Generation`, `NotFound`, `Serialization`, `Validation`, `Duplicate`, `InvalidState`, `Internal`. +- **`Result`** – Alias for `core::result::Result`. + +--- + +## Usage Example + +Here’s a minimal but complete example that builds a site from raw HTML fragments using a custom processor, Tera templates, and the default context builder: + +```rust +use librawssg::{ + Config, ContentRule, Document, FileSystem, Metadata, PipelineBuilder, Processor, + RealFs, RenderContext, Renderer, TeraContextBuilder, TeraRenderer, +}; +use std::path::{Path, PathBuf}; + +// 1. Define a simple processor for .raw files +struct RawProcessor; +impl Processor for RawProcessor { + fn name(&self) -> &'static str { "raw" } + fn can_process(&self, rel: &Path, _orig: &Path) -> bool { + rel.extension().and_then(|e| e.to_str()) == Some("raw") + } + fn process( + &self, + fs: &dyn FileSystem, + rel: &Path, + content_dir: &Path, + ) -> librawssg::Result> { + let full_path = content_dir.join(rel); + let body = fs.read_to_string(&full_path)?; + let title = rel.file_stem().unwrap_or_default().to_string_lossy().to_string(); + let meta = Metadata::new(title, String::new())?; + let url = rel.with_extension("html").to_string_lossy().to_string(); + let output = PathBuf::from(&url); + let doc = Document::new(meta, body, url, output, rel.to_path_buf(), 0, "page".to_string(), false)?; + Ok(Some(doc)) + } +} + +fn main() -> Result<(), Box> { + // 2. Configure the site + let mut config = Config::new().with_site_name("My Site"); + config.add_content_rule(ContentRule::new("page", "**/*.raw", "base.tera")); + config.build.content_dir = "content".to_string(); + config.build.output_dir = "dist".to_string(); + config.build.static_dir = "static".to_string(); + + // 3. Set up the renderer and load templates + let mut renderer = TeraRenderer::new(); + renderer.load_templates_dir(Path::new("templates"))?; + + // 4. Build the pipeline + let pipeline = PipelineBuilder::new() + .config(config) + .content_dir("content") + .output_dir("dist") + .with_fs(Box::new(RealFs)) + .with_renderer(Box::new(renderer)) + .with_context_builder(Box::new(TeraContextBuilder)) + .add_processor(Box::new(RawProcessor)) + .build()?; + + // 5. Run the generation + pipeline.run()?; + println!("Site generated successfully!"); + Ok(()) +} +``` + +For more detailed examples, see the `librawssg_demo` crate in the repository. + +--- + +## Full API Reference + +This section provides a concise reference for every public item re‑exported by `librawssg`. For deeper details, consult the respective sub‑crate documentation (e.g., `librawssg_config`, `librawssg_compiler`). + +### Configuration Types + +All configuration types are in `librawssg::config` (and re‑exported at root). + +- **`Config`** + - `new() -> Self` + - `default() -> Self` + - `with_site_name(self, name: impl Into) -> Self` + - `add_content_rule(&mut self, rule: ContentRule)` + - `find_rule_by_name(&self, name: &str) -> Option<&ContentRule>` + - `remove_rule_by_name(&mut self, name: &str) -> Option` + - `has_duplicate_rule_names(&self) -> bool` + - `validate(&self) -> Result<()>` + - `from_yaml_str(yaml: &str) -> Result` + - `to_yaml_string(&self) -> Result` + - `from_json_str(json: &str) -> Result` + - `to_json_string(&self) -> Result` + +- **`SiteConfig`** + - `new(site_name: impl Into) -> Self` + - `default() -> Self` + - Fields: `navbar: Vec`, `sidebar: Vec`, `site_name: String`, `description: Option`, `language: Option`, `base_url: Option`, `author: Option`, `repo_url: Option`, `license: Option`, `extra: HashMap` + +- **`BuildConfig`** + - `new() -> Self` + - `default() -> Self` + - Fields: `content_dir: String`, `output_dir: String`, `templates_dir: String`, `static_dir: String` + +- **`ContentRule`** + - `new(name: impl Into, pattern: impl Into, template: impl Into) -> Self` + - `default() -> Self` + - Fields: `name: String`, `pattern: String`, `template: String`, `list_template: Option`, `list_enabled: bool`, `extra: HashMap` + +- **`NavItem`** + - `new(label: impl Into, url: impl Into) -> Self` + - `default() -> Self` + - Fields: `label: String`, `url: String`, `children: Vec` + +### Filesystem Types + +- **`FileSystem` trait** (in `librawssg::fs`, re‑exported at root) + - Required methods (see above). + - Provided methods: `is_symlink`, `canonicalize_or_join`, `safe_join`, `copy`, `rename_or_copy`. + +- **`RealFs`** (in `librawssg::fs`, re‑exported at root) + - Implements `FileSystem` using `std::fs`. + - `RealFs` (unit struct), `RealFs::default()`. + +### Handler Types + +- **`Document`** (in `librawssg::handler`, re‑exported at root) + - `new(metadata, body, url, output_path, source_path, depth, content_type, is_list) -> Result` + - `relative_url(&self) -> &str` + - `add_taxonomy(&mut self, name, items)` + - `depth(&self) -> usize` + - `with_list_items(self, items: Vec) -> Self` + - Fields: as listed earlier. + +- **`Metadata`** + - `new(title, description) -> Result` + - `is_draft(&self) -> bool` + - `insert_extra(&mut self, key, value)` + - `get_extra(&self, key: &str) -> Option<&serde_json::Value>` + - Fields: as listed earlier. + +- **`Processor` trait** + - Required: `name`, `can_process`, `process` + - Provided: `priority` + +### Template Types + +- **`RenderContext` trait** + - `as_any(&self) -> &dyn Any` + - `as_mut_any(&mut self) -> &mut dyn Any` + +- **`Renderer` trait** + - `render(&self, template_name: &str, context: &dyn RenderContext) -> Result` + +- **`TeraRenderer`** (feature `tera`, enabled by default) + - `new() -> Self` + - `add_raw_template(&mut self, name: &str, content: &str) -> Result<()>` + - `add_template_file(&mut self, path: &Path) -> Result<()>` + - `add_template_files_from_dir(&mut self, dir: &Path) -> Result<()>` + - `load_templates_dir(&mut self, dir: &Path) -> Result<()>` + - `enable_autoescape(&mut self)` + - `render_str(&self, template_str: &str, context: &dyn RenderContext) -> Result` + - `as_tera(&self) -> &tera::Tera` + - `as_tera_mut(&mut self) -> &mut tera::Tera` + - Implements `Renderer` and `Default`. + +### Compiler Types + +- **`PipelineBuilder`** + - `new() -> Self` + - `config(self, config: Config) -> Self` + - `load_config + Send + Sync>(self, path: P) -> Result` + - `content_dir(self, dir: impl Into) -> Self` + - `output_dir(self, dir: impl Into) -> Self` + - `with_fs(self, fs: Box) -> Self` + - `with_renderer(self, renderer: Box) -> Self` + - `add_processor(self, processor: Box) -> Self` + - `with_context_builder(self, builder: Box) -> Self` + - `add_generator(self, generator: Box) -> Self` + - `build(self) -> Result` + +- **`Pipeline`** + - `run(&self) -> Result<()>` + - `config(&self) -> &Config` + +- **`ContextBuilder` trait** + - `build_context(&self, config: &Config, doc: &Document) -> Result>` + +- **`TeraContextBuilder`** (unit struct) + - Implements `ContextBuilder`. + - `TeraContextBuilder` (no fields), `Default`. + +- **`Generator` trait** + - `generate(&self, pipeline: &Pipeline, output_base: &Path) -> Result<()>` + +- **`match_pattern` function** (accessible via `librawssg::compiler::pattern::match_pattern`) + - Signature: `pub fn match_pattern(pattern: &str, path: &Path) -> bool` + - Supports glob patterns with `*` and `**`. + +### Error Types + +- **`Error` enum** (in `librawssg::error`, re‑exported at root) + - Variants: as listed above. + - Implements `Display`, `Debug`, `std::error::Error`. + - `From` is implemented. + +- **`Result`** type alias + - `pub type Result = core::result::Result` + +--- + +## Feature Flags + +The `tera` feature is enabled by default and provides the `TeraRenderer` implementation. To disable it (e.g., if you use a different template engine), set `default-features = false` in your `Cargo.toml`: + +```toml +[dependencies] +librawssg = { version = "1.0.0", default-features = false } +``` + +Without this feature, the crate still exports the core traits (`Renderer`, `RenderContext`) and all other functionality, but `TeraRenderer` and `TeraContextBuilder` are unavailable. + +--- + +## License + +This project is licensed under the **MIT License**. See the [LICENSE](LICENSE) file for details. + +--- + +_This documentation is generated from the source code of the `librawssg` facade crate and its sub‑crates._ diff --git a/librawssg_compiler/README.md b/librawssg_compiler/README.md index 656345f..6d3ace7 100644 --- a/librawssg_compiler/README.md +++ b/librawssg_compiler/README.md @@ -1,4 +1,759 @@ # librawssg_compiler -Orchestrates the build pipeline for librawssg. -Combines filesystem, processors, renderers, and generators to produce a static site. +**Version**: 1.0.0 (implied) +**Crate name**: `librawssg_compiler` +**Description**: Provides the core compilation pipeline for the `librawssg` static site generator. This crate orchestrates the entire build process: reading content files, processing them via pluggable processors, rendering templates, copying static assets, running custom generators, and outputting the final site atomically. + +--- + +## Table of Contents + +1. [Overview](#overview) +2. [Modules and Re‑exports](#modules-and-re-exports) +3. [`PipelineBuilder`](#struct-pipelinebuilder) + - [Constructor `new()`](#pipelinebuilder-new) + - [Builder Methods](#pipelinebuilder-builder-methods) + - [`config()`](#pipelinebuilder-config) + - [`load_config()`](#pipelinebuilder-load_config) + - [`content_dir()`](#pipelinebuilder-content_dir) + - [`output_dir()`](#pipelinebuilder-output_dir) + - [`with_fs()`](#pipelinebuilder-with_fs) + - [`with_renderer()`](#pipelinebuilder-with_renderer) + - [`add_processor()`](#pipelinebuilder-add_processor) + - [`with_context_builder()`](#pipelinebuilder-with_context_builder) + - [`add_generator()`](#pipelinebuilder-add_generator) + - [`build()` Method](#pipelinebuilder-build) + - [`Default` Implementation](#pipelinebuilder-default) + - [Example Usage](#pipelinebuilder-example) +4. [`ContextBuilder` Trait](#trait-contextbuilder) + - [Required Method `build_context()`](#contextbuilder-build_context) + - [`TeraContextBuilder`](#struct-teracontextbuilder) + - [Implementation Details](#teracontextbuilder-implementation) +5. [`Generator` Trait](#trait-generator) + - [Required Method `generate()`](#generator-generate) +6. [Pattern Matching Function](#function-match_pattern) + - [Signature](#match_pattern-signature) + - [Supported Glob Syntax](#match_pattern-syntax) + - [Algorithm Overview](#match_pattern-algorithm) + - [Examples](#match_pattern-examples) +7. [`Pipeline` Struct](#struct-pipeline) + - [Accessor `config()`](#pipeline-config) + - [Method `run()`](#pipeline-run) + - [Internal Workflow](#pipeline-internal-workflow) + - [Document Processing](#pipeline-document-processing) + - [Content Type Determination](#pipeline-content-type) + - [Rendering](#pipeline-rendering) + - [List Generation](#pipeline-list-generation) + - [Static Assets](#pipeline-static-assets) + - [Generators](#pipeline-generators) + - [Atomic Output Replacement](#pipeline-atomic-output) +8. [Error Handling](#error-handling) +9. [Complete Example from Tests](#complete-example-from-tests) + - [Setting Up a Pipeline](#example-setup) + - [Running the Pipeline](#example-run) + - [Verifying Output](#example-verify) +10. [Testing Suite Overview](#testing-suite-overview) +11. [Conclusion](#conclusion) + +--- + +## Overview + +`librawssg_compiler` is the orchestration layer that ties together all other components of the static site generator: + +- **`Config`** from `librawssg_config` defines content rules and paths. +- **`Processor`** from `librawssg_handler` transforms source files into `Document` objects. +- **`Renderer`** and **`RenderContext`** from `librawssg_templates` handle template rendering. +- **`FileSystem`** from `librawssg_fs` abstracts all I/O operations. +- **`Generator`** (defined here) allows custom post‑processing steps. +- **`ContextBuilder`** (defined here) constructs the render context from a `Document` and `Config`. + +The main entry point is `PipelineBuilder`, which constructs a `Pipeline` after validating the configuration and ensuring mandatory components are present. Running the pipeline performs the full site generation in an atomic fashion, producing output in the configured output directory. + +--- + +## Modules and Re‑exports + +The crate root (`lib.rs`) declares the following public modules: + +- `builder` – Contains `PipelineBuilder`. +- `context` – Contains `ContextBuilder` trait and `TeraContextBuilder`. +- `generator` – Contains `Generator` trait. +- `pattern` – Contains `match_pattern` function. +- `pipeline` – Contains `Pipeline` struct. + +Re‑exported types at the crate root: + +```rust +pub use builder::PipelineBuilder; +pub use context::ContextBuilder; +pub use context::TeraContextBuilder; +pub use generator::Generator; +pub use pipeline::Pipeline; +``` + +The `match_pattern` function is also re‑exported? Actually it is in `pattern` module and not re‑exported at root, so users must use `librawssg_compiler::pattern::match_pattern`. However, in `pipeline.rs` it is imported via `crate::pattern::match_pattern`, but for external users they need to access it via module path. + +--- + +## Struct `PipelineBuilder` + +`PipelineBuilder` is a builder‑style struct that collects all components needed to run the compilation pipeline and then builds a `Pipeline`. + +```rust +pub struct PipelineBuilder { + config: Config, + content_dir: PathBuf, + output_dir: PathBuf, + fs: Box, + renderer: Option>, + processors: Vec>, + context_builder: Option>, + generators: Vec>, +} +``` + +**Note**: Fields are private; use the builder methods to configure. + +### `PipelineBuilder::new` + +```rust +#[must_use] +pub fn new() -> Self +``` + +**Purpose**: Creates a new builder with default values: + +- `config`: `Config::default()` +- `content_dir`: `"content"` +- `output_dir`: `"dist"` +- `fs`: `Box::new(RealFs)` +- `renderer`: `None` +- `processors`: empty vector +- `context_builder`: `None` +- `generators`: empty vector + +**Returns**: A fresh `PipelineBuilder`. + +**Example**: + +```rust +let builder = PipelineBuilder::new(); +``` + +--- + +### PipelineBuilder Builder Methods + +All builder methods consume `self` and return `Self`, allowing method chaining. + +#### `config` + +```rust +#[must_use] +pub fn config(mut self, config: Config) -> Self +``` + +**Purpose**: Sets the configuration object. + +**Parameters**: + +- `config`: A `Config` instance from `librawssg_config`. + +**Returns**: The builder with the config set. + +#### `load_config` + +```rust +pub fn load_config + Send + Sync>(mut self, path: P) -> Result +``` + +**Purpose**: Reads a YAML config file from disk and parses it into a `Config`. Errors are converted to `Error::Config`. + +**Parameters**: + +- `path`: Path to the YAML file. + +**Returns**: `Ok(Self)` if parsing succeeds, otherwise `Err(Error::Config)`. + +**Note**: The file is read using standard `std::fs::read_to_string`; the error is wrapped in `Error::Config`. + +#### `content_dir` + +```rust +#[must_use] +pub fn content_dir(mut self, dir: impl Into) -> Self +``` + +**Purpose**: Overrides the content directory. + +**Parameters**: + +- `dir`: Any type convertible to `PathBuf`. + +**Default**: `"content"` (but may be overridden by config if not explicitly set; see `build()`). + +#### `output_dir` + +```rust +#[must_use] +pub fn output_dir(mut self, dir: impl Into) -> Self +``` + +**Purpose**: Overrides the output directory. + +**Default**: `"dist"`. + +#### `with_fs` + +```rust +#[must_use] +pub fn with_fs(mut self, fs: Box) -> Self +``` + +**Purpose**: Sets a custom filesystem implementation. Useful for testing or non‑standard backends. + +**Default**: `RealFs`. + +#### `with_renderer` + +```rust +#[must_use] +pub fn with_renderer(mut self, renderer: Box) -> Self +``` + +**Purpose**: Sets the template renderer. **Mandatory**; `build()` will fail if not set. + +#### `add_processor` + +```rust +#[must_use] +pub fn add_processor(mut self, processor: Box) -> Self +``` + +**Purpose**: Adds a content processor to the pipeline. Multiple processors can be added; they are tried in order for each source file. + +#### `with_context_builder` + +```rust +#[must_use] +pub fn with_context_builder(mut self, builder: Box) -> Self +``` + +**Purpose**: Sets the context builder. **Mandatory**; `build()` will fail if not set. + +#### `add_generator` + +```rust +#[must_use] +pub fn add_generator(mut self, generator: Box) -> Self +``` + +**Purpose**: Adds a post‑processing generator. Generators run after all documents are rendered and static assets copied. + +--- + +### PipelineBuilder::build + +```rust +pub fn build(mut self) -> Result +``` + +**Purpose**: Validates the configuration, ensures required components are present, and constructs a `Pipeline`. + +**Behavior**: + +1. Calls `self.config.validate()?` (see `librawssg_config::Config::validate`). +2. If `content_dir` is still the default `"content"` (i.e., not changed by `content_dir()`), it is replaced with `self.config.build.content_dir`. +3. Similarly, if `output_dir` is still `"dist"`, it is replaced with `self.config.build.output_dir`. +4. Takes the renderer from `self.renderer` (using `take()`). If `None`, returns `Error::Config("template renderer not set")`. +5. Takes the context builder from `self.context_builder`. If `None`, returns `Error::Config("context builder not set")`. +6. Moves all remaining fields into a new `Pipeline` and returns `Ok`. + +**Returns**: + +- `Ok(Pipeline)` on success. +- `Err(Error::Validation)` if config invalid. +- `Err(Error::Config)` if renderer or context builder missing. + +--- + +### PipelineBuilder Default + +```rust +impl Default for PipelineBuilder { + fn default() -> Self { + Self::new() + } +} +``` + +Allows creating with `PipelineBuilder::default()`. + +--- + +### PipelineBuilder Example + +```rust +use librawssg_compiler::{PipelineBuilder, TeraContextBuilder}; +use librawssg_config::Config; +use librawssg_fs::RealFs; +use librawssg_handler::Processor; +use librawssg_templates::{Renderer, TeraRenderer}; + +// Assuming custom processor and renderer exist +let config = Config::new().with_site_name("My Site"); +let processor = Box::new(MyProcessor); +let renderer = Box::new(TeraRenderer::new()); // TeraRenderer must be configured with templates beforehand +let context_builder = Box::new(TeraContextBuilder); + +let pipeline = PipelineBuilder::new() + .config(config) + .content_dir("src") + .output_dir("public") + .with_fs(Box::new(RealFs)) + .with_renderer(renderer) + .with_context_builder(context_builder) + .add_processor(processor) + .build()?; +``` + +--- + +## Trait `ContextBuilder` + +```rust +pub trait ContextBuilder: Send + Sync { + fn build_context(&self, config: &Config, doc: &Document) -> Result>; +} +``` + +**Purpose**: Abstract factory that creates a `RenderContext` from the global `Config` and a specific `Document`. The renderer then uses this context to render the document's template. + +**Requirements**: Implementors must be `Send + Sync`. + +### `build_context` + +**Parameters**: + +- `config`: Reference to the site configuration. +- `doc`: Reference to the document being rendered. + +**Returns**: + +- `Ok(Box)` – a boxed trait object holding the render context. +- `Err(librawssg_error::Error)` if context creation fails. + +--- + +## Struct `TeraContextBuilder` + +```rust +#[derive(Debug, Default, Clone, Copy)] +pub struct TeraContextBuilder; +``` + +A concrete implementation of `ContextBuilder` that produces a `tera::Context` populated with common page and site data. + +### Implementation Details + +`TeraContextBuilder` creates a new `tera::Context` and inserts the following keys: + +| Key | Value Source | Description | +| ------------------ | --------------------------- | --------------------------------------------- | +| `site` | `&config.site` | The full `SiteConfig` object. | +| `page_title` | `&doc.metadata.title` | Document title. | +| `page_description` | `&doc.metadata.description` | Document description. | +| `page_author` | `&doc.metadata.author` | Optional author. | +| `page_date` | `&doc.metadata.date` | Optional publication date. | +| `page_tags` | `&doc.metadata.tags` | Vector of tags. | +| `page_content` | `&doc.body` | The rendered body content of the document. | +| `page_url` | `&doc.url` | Relative URL of the document. | +| `page_depth` | `&doc.depth` | Depth in the site hierarchy. | +| `page_type` | `&doc.content_type` | Content type identifier (e.g., `"blog"`). | +| `page_is_list` | `&doc.is_list` | Boolean indicating list page. | +| `page_list_items` | `&doc.list_items` | Optional vector of child documents for lists. | + +**Note**: The `page_list_items` field is inserted as `&doc.list_items` which is `Option>`. In Tera templates, this will be `None` or an array. + +**Example**: + +```rust +use librawssg_compiler::TeraContextBuilder; +let builder = TeraContextBuilder; +let ctx = builder.build_context(&config, &doc)?; +// Pass `ctx` to renderer.render(...) +``` + +--- + +## Trait `Generator` + +```rust +pub trait Generator: Send + Sync { + fn generate(&self, pipeline: &Pipeline, output_base: &Path) -> Result<()>; +} +``` + +**Purpose**: Allows custom post‑processing steps after the main site generation. Generators can write additional files to the output directory (e.g., RSS feed, sitemap, search index). + +**Requirements**: Implementors must be `Send + Sync`. + +### `generate` + +**Parameters**: + +- `pipeline`: Reference to the running `Pipeline`, which provides access to its configuration, filesystem, etc. (though the fields are crate‑private, the `config()` method is available). +- `output_base`: Path to the temporary output directory where generated files should be written. The pipeline's `run()` method later moves this directory atomically to the final output location. + +**Returns**: + +- `Ok(())` on success. +- `Err(librawssg_error::Error)` on failure. + +**Example** (from tests): + +```rust +struct DummyGenerator; + +impl Generator for DummyGenerator { + fn generate(&self, _pipeline: &Pipeline, output_base: &Path) -> Result<()> { + let path = output_base.join("generated.txt"); + std::fs::write(&path, b"generated")?; + Ok(()) + } +} +``` + +--- + +## Function `match_pattern` + +```rust +#[must_use] +pub fn match_pattern(pattern: &str, path: &Path) -> bool +``` + +**Location**: `librawssg_compiler::pattern` + +**Purpose**: Checks whether a file path matches a glob pattern with support for `*` (within a segment) and `**` (across segments). Used by the pipeline to assign content types based on `ContentRule` patterns. + +**Parameters**: + +- `pattern`: A glob‑like pattern string, e.g., `"blog/**/*.html"`. +- `path`: A `Path` to test (usually a relative path from the content directory). + +**Returns**: `true` if the path matches the pattern; `false` otherwise. + +### Supported Glob Syntax + +- `*` – Matches any sequence of characters within a single path segment (i.e., does not cross `/`). +- `**` – Matches any number of path segments, including zero. + +**Limitations**: + +- Only `*` and `**` are supported; no character classes (`[abc]`) or alternation. +- Patterns are split on `/`; backslashes are not treated as separators (path normalization may be needed on Windows). +- The implementation uses a custom recursive algorithm; for complex patterns, behavior may differ from standard glob libraries. + +### Algorithm Overview + +The function first converts the path to a string (lossy) and splits both pattern and path by `/`. It then calls an internal recursive `match_pattern_slice`. The logic: + +- If both pattern and path segments are exhausted → `true`. +- If pattern still has segments but path is empty → only `**` segments are allowed. +- If first pattern segment is `"**"`: + - If it's the only segment → `true`. + - Otherwise, try matching the rest of the pattern against every suffix of the path. +- Otherwise, the first pattern segment must match the first path segment using `segment_matches` (handles `*` wildcards), and recursion continues on the rest. +- `segment_matches` handles `*` by trying to match the remainder of the pattern segment against suffixes of the path segment. + +### Examples + +```rust +use std::path::Path; +use librawssg_compiler::pattern::match_pattern; + +assert!(match_pattern("**/*.html", Path::new("blog/post.html"))); +assert!(match_pattern("blog/**/*.html", Path::new("blog/2024/post.html"))); +assert!(match_pattern("*.html", Path::new("index.html"))); +assert!(!match_pattern("*.html", Path::new("blog/post.html"))); +assert!(match_pattern("**", Path::new("anything/at/all"))); +assert!(match_pattern("**/*.md", Path::new("readme.md"))); // zero segments before .md +``` + +--- + +## Struct `Pipeline` + +The `Pipeline` is the core execution engine. It is created by `PipelineBuilder::build()` and holds all necessary components. + +```rust +pub struct Pipeline { + config: Config, + fs: Box, + renderer: Box, + processors: Vec>, + context_builder: Box, + generators: Vec>, + content_dir: PathBuf, + output_dir: PathBuf, +} +``` + +All fields are private; access to configuration is provided via the `config()` method. + +### Pipeline::config + +```rust +#[must_use] +pub const fn config(&self) -> &Config +``` + +**Purpose**: Returns a reference to the configuration used by this pipeline. + +**Returns**: `&Config`. + +--- + +### Pipeline::run + +```rust +pub fn run(&self) -> Result<()> +``` + +**Purpose**: Executes the full site generation process atomically. + +**Behavior**: + +1. Determines a temporary output directory: `output_dir.with_extension("tmp")`. For example, if `output_dir` is `"dist"`, the temp dir is `"dist.tmp"`. +2. If the temp dir already exists, it is removed. +3. Creates the temp dir. +4. Calls internal `generate_to(&tmp_dir)` to perform the actual generation into the temporary location. +5. If the final output directory exists, it is removed. +6. Attempts to rename the temp dir to the final output dir. + - On success, returns `Ok(())`. + - If rename fails with `CrossesDevices` error (different filesystems), fallback: copy the temp dir contents to the final output dir using `copy_dir_all` (internal method), then remove the temp dir. + - For any other error, returns `Error::Generation("atomic rename failed: ...")`. + +**Returns**: `Ok(())` on success, or `Err(Error)`. + +**Note**: This atomic approach ensures that the final output directory is never left in a partially generated state; either the old output remains untouched (if generation fails) or the new output replaces it atomically (or near‑atomically). + +--- + +### Pipeline Internal Workflow + +The internal method `generate_to(output_base: &Path)` orchestrates the entire generation. The following steps are performed (not public, but described for understanding): + +1. **Create output directory** – `fs.create_dir_all(output_base)`. +2. **Process documents** – Calls `process_documents()` to get a vector of `Document`. +3. **Group documents by content type** – Builds a `HashMap>`. +4. **Render non‑list documents** – For each `Document` where `is_list == false`, calls `render_document(output_base, doc)`. +5. **Generate list pages** – For each content type that has a matching `ContentRule` with `list_enabled == true`, `list_template` set, and at least one document, creates a synthetic list document (`is_list = true`) and renders it using the list template. The list document’s `list_items` are all documents of that content type. The output path is `"{content_type}/index.html"`. +6. **Copy static assets** – If the directory specified by `config.build.static_dir` exists, its entire contents are copied to `output_base/static_dir_name` (using `copy_dir_all`). +7. **Run generators** – Iterates over all `generators` and calls `generate()` for each, passing the output base. + +#### Document Processing + +`process_documents()` walks the content directory (using `fs.walk_dir`). For each file, it: + +- Computes the path relative to `content_dir`. +- Iterates through the processors in order; the first processor whose `can_process(rel, &file_path)` returns `true` is used. +- Calls that processor’s `process(&*fs, rel, &content_dir)`, which returns `Option`. +- If a document is returned, its `content_type` is overridden based on the matching content rule (via `determine_content_type`), and its `depth` is set to the number of path components minus 1. +- The document is added to the result list. + +If no processor matches, the file is ignored. If a processor returns `Ok(None)`, the file is skipped. If any processor returns an error, the whole processing fails. + +#### Content Type Determination + +`determine_content_type(rel)` iterates through the `config.content_rules` in **reverse order** (so later rules take precedence) and returns the `name` of the first rule whose `pattern` matches the relative path. If no rule matches, the content type defaults to `"page"`. + +#### Rendering + +`render_document` calls `template_for_document(doc)` to get the template name: + +- Iterates through content rules, finds the rule whose `name` equals `doc.content_type`. +- If `doc.is_list` and the rule has a `list_template`, that template is used. +- Otherwise, the rule’s `template` is used. +- If no rule is found, returns `Error::Generation("no content rule found for type '...'")`. + +The actual rendering uses: + +- `context_builder.build_context(&config, doc)` to get a `RenderContext`. +- `renderer.render(template, &*ctx)` to produce the HTML string. +- `write_output(output_base, doc, html.as_bytes())` to write the result. + +#### List Generation + +When generating list pages, the pipeline creates a `Metadata` with `title = content_type` and empty description. Then constructs a `Document` with: + +- `body`: empty string (the list template is expected to use `page_list_items`). +- `url`: `"{content_type}/index.html"`. +- `output_path`: `"{content_type}/index.html"`. +- `source_path`: `PathBuf::from("__list__")` (placeholder). +- `depth`: `1`. +- `content_type`: same as the content type. +- `is_list`: `true`. + +Then sets `list_items` to the cloned vector of documents of that type, and renders using the list template. + +#### Static Assets + +The static directory from `config.build.static_dir` is copied verbatim. The destination is `output_base` joined with the static directory’s base name (e.g., if `static_dir = "static"`, files are copied to `output_base/static/`). Directories are recursively copied. + +#### Generators + +After all documents and static files are in place, each registered `Generator` is invoked with `&self` and `output_base`. This allows adding custom files like RSS feeds, sitemaps, etc. + +--- + +## Error Handling + +`librawssg_compiler` uses `librawssg_error::Error` for all fallible operations. The common error variants encountered: + +- `Error::Config` – For configuration issues (e.g., missing renderer, invalid config file). +- `Error::Validation` – From `Config::validate`. +- `Error::Generation` – For errors during pipeline execution (e.g., write failures, unsafe output path). +- `Error::Io` – From filesystem operations (though these may be wrapped in `Error::Generation` in some places). +- `Error::Render` – From template rendering. + +Methods return `Result` (alias for `std::result::Result`). + +--- + +## Complete Example from Tests + +The test file `full_compiler_test.rs` demonstrates a full working pipeline with mock components. Below is a simplified but complete example. + +### Example Setup + +Define mock renderer, context, context builder, and processor: + +```rust +use librawssg_compiler::{PipelineBuilder, TeraContextBuilder}; +use librawssg_config::{Config, ContentRule}; +use librawssg_fs::{FileSystem, RealFs}; +use librawssg_handler::{Document, Metadata, Processor}; +use librawssg_templates::{RenderContext, Renderer}; +use std::path::{Path, PathBuf}; + +// Mock renderer: returns "rendered:{template_name}" +struct MockRenderer; +impl Renderer for MockRenderer { + fn render(&self, template_name: &str, _ctx: &dyn RenderContext) -> Result { + Ok(format!("rendered:{template_name}")) + } +} + +// Mock context +struct MockContext; +impl RenderContext for MockContext { + fn as_any(&self) -> &dyn Any { self } + fn as_mut_any(&mut self) -> &mut dyn Any { self } +} + +// Mock context builder +struct MockContextBuilder; +impl ContextBuilder for MockContextBuilder { + fn build_context(&self, _config: &Config, _doc: &Document) -> Result> { + Ok(Box::new(MockContext)) + } +} + +// Processor that handles .html files +struct RawHtmlProcessor; +impl Processor for RawHtmlProcessor { + fn name(&self) -> &'static str { "raw-html" } + fn can_process(&self, relative_path: &Path, _original_path: &Path) -> bool { + relative_path.extension().is_some_and(|ext| ext == "html") + } + fn process(&self, fs: &dyn FileSystem, relative_path: &Path, content_dir: &Path) -> Result> { + let full_path = content_dir.join(relative_path); + let content = fs.read_to_string(&full_path)?; + let url = relative_path.with_extension("html").to_string_lossy().to_string(); + let output_path = PathBuf::from(&url); + let metadata = Metadata::new("Test", "Description")?; + let doc = Document::new( + metadata, + content, + url, + output_path, + relative_path.to_path_buf(), + 0, + "page".to_string(), + false, + )?; + Ok(Some(doc)) + } +} + +// Config with one rule +fn setup_config() -> Config { + let mut config = Config::new().with_site_name("Compiler Test"); + config.add_content_rule(ContentRule::new("page", "**/*.html", "base")); + config +} + +// Build pipeline +fn build_pipeline(content_dir: &Path, output_dir: &Path) -> Result { + PipelineBuilder::new() + .config(setup_config()) + .content_dir(content_dir) + .output_dir(output_dir) + .with_fs(Box::new(RealFs)) + .with_renderer(Box::new(MockRenderer)) + .with_context_builder(Box::new(MockContextBuilder)) + .add_processor(Box::new(RawHtmlProcessor)) + .build() +} +``` + +### Running the Pipeline + +```rust +let tmp = tempfile::TempDir::new()?; +let content_dir = tmp.path().join("content"); +let output_dir = tmp.path().join("dist"); +std::fs::create_dir_all(&content_dir)?; +std::fs::write(content_dir.join("index.html"), "

Home

")?; + +let pipeline = build_pipeline(&content_dir, &output_dir)?; +pipeline.run()?; +``` + +### Verifying Output + +```rust +let output_file = output_dir.join("index.html"); +assert!(output_file.exists()); +let content = std::fs::read_to_string(output_file)?; +assert_eq!(content, "rendered:base"); +``` + +--- + +## Testing Suite Overview + +The test file `tests/full_compiler_test.rs` contains comprehensive integration tests covering: + +- Generation of a single page. +- Content type rules (different templates for different content types). +- List page generation. +- Static asset copying. +- Custom generators. +- Atomic replacement of existing output directory. +- Handling empty content directory. +- Skipping files that no processor handles. +- Error cases: missing renderer, missing context builder, invalid config. + +Each test uses temporary directories (`tempfile`) and mock implementations to isolate components. The tests serve as executable examples of how to configure and run the pipeline. + +--- + +## Conclusion + +`librawssg_compiler` is the central execution engine of the static site generator. It provides a flexible builder to assemble the necessary components, a robust `Pipeline` that orchestrates all steps, and extension points via `Processor`, `Renderer`, `ContextBuilder`, and `Generator`. The pattern matching function and atomic output generation ensure correctness and safety. The comprehensive test suite demonstrates practical usage and edge cases. + +For further details, refer to the source code and the test file. diff --git a/librawssg_config/README.md b/librawssg_config/README.md index 466003a..c48d4ea 100644 --- a/librawssg_config/README.md +++ b/librawssg_config/README.md @@ -1,4 +1,762 @@ # librawssg_config -Configuration types and validation for librawssg. -Provides `Config`, site metadata, build settings, and content rule definitions. +**Version**: 1.0.0 (implied) +**Crate name**: `librawssg_config` +**Description**: Defines configuration data structures for the `librawssg` static site generator. Provides `Config`, `SiteConfig`, `BuildConfig`, `ContentRule`, and `NavItem` types, along with serialization/deserialization support (YAML and JSON) and validation logic. + +--- + +## Table of Contents + +1. [Overview](#overview) +2. [Modules](#modules) +3. [`BuildConfig`](#struct-buildconfig) + - [Fields](#buildconfig-fields) + - [Constructor `new()`](#buildconfig-new) + - [`Default` Implementation](#buildconfig-default) + - [Serialization & Deserialization](#buildconfig-serialization) +4. [`ContentRule`](#struct-contentrule) + - [Fields](#contentrule-fields) + - [Constructor `new()`](#contentrule-new) + - [`Default` Implementation](#contentrule-default) + - [Serialization & Deserialization](#contentrule-serialization) +5. [`NavItem`](#struct-navitem) + - [Fields](#navitem-fields) + - [Constructor `new()`](#navitem-new) + - [`Default` Implementation](#navitem-default) + - [Serialization & Deserialization](#navitem-serialization) +6. [`SiteConfig`](#struct-siteconfig) + - [Fields](#siteconfig-fields) + - [Constructor `new()`](#siteconfig-new) + - [`Default` Implementation](#siteconfig-default) + - [Serialization & Deserialization](#siteconfig-serialization) +7. [`Config`](#struct-config) + - [Fields](#config-fields) + - [Constructor `new()`](#config-new) + - [Builder Method `with_site_name()`](#config-with_site_name) + - [Rule Management Methods](#config-rule-management) + - [`add_content_rule()`](#config-add_content_rule) + - [`find_rule_by_name()`](#config-find_rule_by_name) + - [`remove_rule_by_name()`](#config-remove_rule_by_name) + - [`has_duplicate_rule_names()`](#config-has_duplicate_rule_names) + - [Validation Method `validate()`](#config-validate) + - [Serialization Methods](#config-serialization) + - [`from_yaml_str()`](#config-from_yaml_str) + - [`to_yaml_string()`](#config-to_yaml_string) + - [`from_json_str()`](#config-from_json_str) + - [`to_json_string()`](#config-to_json_string) +8. [Error Handling](#error-handling) +9. [Serialization Details](#serialization-details) +10. [Examples from Tests](#examples-from-tests) +11. [Testing Suite Overview](#testing-suite-overview) +12. [Conclusion](#conclusion) + +--- + +## Overview + +`librawssg_config` provides the central configuration types used by the static site generator. The main `Config` struct combines site settings (`SiteConfig`), build paths (`BuildConfig`), content processing rules (`ContentRule`), and arbitrary extra data. All types are serializable/deserializable via `serde`, enabling configuration to be read from and written to YAML or JSON files. + +The types are designed with sensible defaults and include a `validate()` method to ensure the configuration is internally consistent and safe (e.g., preventing path traversal in patterns). + +--- + +## Modules + +The crate root (`lib.rs`) declares the following public modules: + +- `build` – Contains `BuildConfig`. +- `config` – Contains `Config`. +- `content_rule` – Contains `ContentRule`. +- `nav` – Contains `NavItem`. +- `site` – Contains `SiteConfig`. + +All public types are re‑exported at the crate root for convenience: + +```rust +pub use build::BuildConfig; +pub use config::Config; +pub use content_rule::ContentRule; +pub use nav::NavItem; +pub use site::SiteConfig; +``` + +--- + +## Struct `BuildConfig` + +Represents filesystem path configuration for the build process. + +```rust +#[derive(Debug, Clone, PartialEq, Eq, Serialize, Deserialize)] +#[non_exhaustive] +pub struct BuildConfig { + #[serde(default = "default_content_dir")] + pub content_dir: String, + #[serde(default = "default_output_dir")] + pub output_dir: String, + #[serde(default = "default_templates_dir")] + pub templates_dir: String, + #[serde(default = "default_static_dir")] + pub static_dir: String, +} +``` + +### BuildConfig Fields + +| Field | Type | Default Value | Description | +| --------------- | -------- | ------------- | ------------------------------------------------------ | +| `content_dir` | `String` | `"content"` | Directory containing source content files. | +| `output_dir` | `String` | `"dist"` | Directory where generated site output will be written. | +| `templates_dir` | `String` | `"templates"` | Directory containing template files. | +| `static_dir` | `String` | `"static"` | Directory containing static assets (copied as-is). | + +**Note**: `#[non_exhaustive]` prevents external crates from exhaustively matching or constructing with a struct literal. Use the provided constructors or update syntax. + +### BuildConfig::new + +```rust +#[must_use] +pub fn new() -> Self +``` + +**Purpose**: Creates a `BuildConfig` with default values (identical to `BuildConfig::default()`). + +**Returns**: A new `BuildConfig` with all fields set to their defaults. + +**Example**: + +```rust +let build = BuildConfig::new(); +assert_eq!(build.content_dir, "content"); +``` + +### BuildConfig Default + +The `Default` trait is implemented with the following values: + +- `content_dir`: `"content"` +- `output_dir`: `"dist"` +- `templates_dir`: `"templates"` +- `static_dir`: `"static"` + +These defaults can be overridden during deserialization; missing fields in serialized data will fall back to these defaults (thanks to `#[serde(default = "...")]`). + +### BuildConfig Serialization + +`BuildConfig` derives `Serialize` and `Deserialize`. When deserializing from YAML/JSON, any omitted fields will use the specified default functions. This allows partial configuration. + +**Example** (from tests): + +```rust +let yaml = "content_dir: custom_content\noutput_dir: public\n"; +let build: BuildConfig = serde_yaml::from_str(yaml)?; +assert_eq!(build.content_dir, "custom_content"); +assert_eq!(build.templates_dir, "templates"); // default +``` + +--- + +## Struct `ContentRule` + +Defines how a certain group of content files should be processed. + +```rust +#[derive(Debug, Clone, PartialEq, Eq, Serialize, Deserialize, Default)] +#[non_exhaustive] +pub struct ContentRule { + pub name: String, + pub pattern: String, + pub template: String, + #[serde(default)] + pub list_template: Option, + #[serde(default)] + pub list_enabled: bool, + #[serde(default)] + pub extra: HashMap, +} +``` + +### ContentRule Fields + +| Field | Type | Default | Description | +| --------------- | ------------------------------------ | --------- | ------------------------------------------------------------------------------- | +| `name` | `String` | `""` | Unique identifier for the rule (e.g., `"blog"`, `"page"`). | +| `pattern` | `String` | `""` | Glob pattern matching content files (e.g., `"**/*.md"`). Must not contain `..`. | +| `template` | `String` | `""` | Name of the template to use for rendering each matched file. | +| `list_template` | `Option` | `None` | Optional template name for rendering list pages (e.g., index pages). | +| `list_enabled` | `bool` | `false` | Whether list generation is enabled for this rule. | +| `extra` | `HashMap` | empty map | Arbitrary extra data associated with the rule. | + +### ContentRule::new + +```rust +#[must_use] +pub fn new( + name: impl Into, + pattern: impl Into, + template: impl Into, +) -> Self +``` + +**Purpose**: Creates a `ContentRule` with the required fields (`name`, `pattern`, `template`). All other fields are set to their defaults. + +**Parameters**: + +- `name`: The rule name (converted to `String`). +- `pattern`: The glob pattern (converted to `String`). +- `template`: The template name (converted to `String`). + +**Returns**: A new `ContentRule` instance. + +**Example**: + +```rust +let rule = ContentRule::new("blog", "**/*.md", "post"); +assert_eq!(rule.name, "blog"); +assert!(!rule.list_enabled); +``` + +### ContentRule Default + +The `Default` implementation (derived) sets: + +- `name`, `pattern`, `template`: empty strings +- `list_template`: `None` +- `list_enabled`: `false` +- `extra`: empty map + +### ContentRule Serialization + +Serializes/deserializes with `serde`. Missing optional fields default as specified. The `extra` map can hold any JSON‑compatible values. + +**Example**: + +```rust +let mut rule = ContentRule::new("page", "**/*.html", "base"); +rule.list_enabled = true; +rule.list_template = Some("list".into()); +rule.extra.insert("key".into(), json!("value")); +let yaml = serde_yaml::to_string(&rule)?; +let parsed: ContentRule = serde_yaml::from_str(&yaml)?; +assert_eq!(rule, parsed); +``` + +--- + +## Struct `NavItem` + +Represents an item in a navigation menu (navbar or sidebar). + +```rust +#[derive(Debug, Clone, PartialEq, Eq, Serialize, Deserialize, Default)] +#[non_exhaustive] +pub struct NavItem { + pub label: String, + pub url: String, + pub children: Vec, +} +``` + +### NavItem Fields + +| Field | Type | Default | Description | +| ---------- | -------------- | ------- | ---------------------------------------------- | +| `label` | `String` | `""` | Display text for the navigation link. | +| `url` | `String` | `""` | URL the link points to (relative or absolute). | +| `children` | `Vec` | empty | Nested sub‑items, enabling hierarchical menus. | + +### NavItem::new + +```rust +#[must_use] +pub fn new(label: impl Into, url: impl Into) -> Self +``` + +**Purpose**: Creates a `NavItem` with a label and URL. The `children` vector starts empty. + +**Parameters**: + +- `label`: Display label. +- `url`: Target URL. + +**Returns**: A new `NavItem`. + +**Example**: + +```rust +let item = NavItem::new("Home", "/"); +assert_eq!(item.label, "Home"); +assert!(item.children.is_empty()); +``` + +### NavItem Default + +`NavItem::default()` creates an item with empty label, empty URL, and no children. + +### NavItem Serialization + +Supports serialization and deserialization via `serde`. Nested children are handled recursively. + +**Example**: + +```rust +let parent = NavItem::new("Docs", "/docs"); +parent.children.push(NavItem::new("API", "/docs/api")); +let json = serde_json::to_string(&parent)?; +let parsed: NavItem = serde_json::from_str(&json)?; +assert_eq!(parent, parsed); +``` + +--- + +## Struct `SiteConfig` + +Holds global site metadata and navigation structures. + +```rust +#[derive(Debug, Clone, PartialEq, Eq, Serialize, Deserialize)] +#[non_exhaustive] +pub struct SiteConfig { + #[serde(default)] + pub navbar: Vec, + #[serde(default)] + pub sidebar: Vec, + #[serde(default = "default_site_name")] + pub site_name: String, + #[serde(default)] + pub description: Option, + #[serde(default = "default_language")] + pub language: Option, + #[serde(default)] + pub base_url: Option, + #[serde(default)] + pub author: Option, + #[serde(default)] + pub repo_url: Option, + #[serde(default)] + pub license: Option, + #[serde(default)] + pub extra: HashMap, +} +``` + +### SiteConfig Fields + +| Field | Type | Default | Description | +| ------------- | ------------------------------------ | ------------- | ---------------------------------------------------------------------- | +| `navbar` | `Vec` | empty | Navigation items for the top bar. | +| `sidebar` | `Vec` | empty | Navigation items for the sidebar. | +| `site_name` | `String` | `"librawssg"` | The name of the website. | +| `description` | `Option` | `None` | Short site description. | +| `language` | `Option` | `Some("en")` | Site language code (e.g., `"en"`, `"id"`). | +| `base_url` | `Option` | `None` | Base URL for the site; must start with `http://` or `https://` if set. | +| `author` | `Option` | `None` | Default author name. | +| `repo_url` | `Option` | `None` | URL to the source repository. | +| `license` | `Option` | `None` | License identifier (e.g., `"MIT"`). | +| `extra` | `HashMap` | empty map | Arbitrary extra site‑wide metadata. | + +### SiteConfig::new + +```rust +#[must_use] +pub fn new(site_name: impl Into) -> Self +``` + +**Purpose**: Creates a `SiteConfig` with a custom site name. All other fields are set to their defaults (navbar/sidebar empty, language `Some("en")`, etc.). + +**Parameters**: + +- `site_name`: The site name (converted to `String`). + +**Returns**: A new `SiteConfig`. + +**Example**: + +```rust +let site = SiteConfig::new("My Site"); +assert_eq!(site.site_name, "My Site"); +assert_eq!(site.language.as_deref(), Some("en")); +``` + +### SiteConfig Default + +`SiteConfig::default()` sets: + +- `site_name`: `"librawssg"` +- `language`: `Some("en")` +- All `Option` fields: `None` +- Vectors and map: empty + +### SiteConfig Serialization + +Supports YAML/JSON. Missing fields during deserialization use defaults. The `language` default is provided by a custom function. + +**Example**: + +```rust +let mut site = SiteConfig::new("Test"); +site.extra.insert("foo".into(), json!("bar")); +let yaml = serde_yaml::to_string(&site)?; +let parsed: SiteConfig = serde_yaml::from_str(&yaml)?; +assert_eq!(site, parsed); +``` + +--- + +## Struct `Config` + +The top‑level configuration combining all other components. + +```rust +#[derive(Debug, Clone, PartialEq, Eq, Serialize, Deserialize, Default)] +#[non_exhaustive] +pub struct Config { + pub site: SiteConfig, + pub build: BuildConfig, + pub content_rules: Vec, + #[serde(default)] + pub extra: HashMap, +} +``` + +### Config Fields + +| Field | Type | Default | Description | +| --------------- | ------------------------------------ | ----------------------------------------------- | --------------------------------- | +| `site` | `SiteConfig` | `SiteConfig::default()` (site name "librawssg") | Global site configuration. | +| `build` | `BuildConfig` | `BuildConfig::default()` | Build path settings. | +| `content_rules` | `Vec` | empty | List of content processing rules. | +| `extra` | `HashMap` | empty map | Arbitrary top‑level extra data. | + +### Config::new + +```rust +#[must_use] +pub fn new() -> Self +``` + +**Purpose**: Creates a `Config` with all fields defaulted. Equivalent to `Config::default()`. + +**Returns**: A new `Config`. + +**Example**: + +```rust +let config = Config::new(); +assert!(config.content_rules.is_empty()); +``` + +### Config::with_site_name + +```rust +#[must_use] +pub fn with_site_name(mut self, name: impl Into) -> Self +``` + +**Purpose**: Builder‑style method that sets the `site.site_name` and returns the modified `Config`. + +**Parameters**: + +- `name`: The new site name. + +**Returns**: The same `Config` with updated site name. + +**Example**: + +```rust +let config = Config::new().with_site_name("My Awesome Site"); +assert_eq!(config.site.site_name, "My Awesome Site"); +``` + +### Config Rule Management + +#### `add_content_rule` + +```rust +pub fn add_content_rule(&mut self, rule: ContentRule) +``` + +**Purpose**: Appends a `ContentRule` to the `content_rules` vector. + +**Parameters**: + +- `rule`: The rule to add. + +**Example**: + +```rust +config.add_content_rule(ContentRule::new("blog", "**/*.md", "post")); +``` + +#### `find_rule_by_name` + +```rust +#[must_use] +pub fn find_rule_by_name(&self, name: &str) -> Option<&ContentRule> +``` + +**Purpose**: Searches for a content rule by its `name` field. + +**Parameters**: + +- `name`: The rule name to search for. + +**Returns**: `Some(&ContentRule)` if found, otherwise `None`. + +**Example**: + +```rust +if let Some(rule) = config.find_rule_by_name("blog") { + // ... +} +``` + +#### `remove_rule_by_name` + +```rust +pub fn remove_rule_by_name(&mut self, name: &str) -> Option +``` + +**Purpose**: Removes and returns the first content rule whose `name` matches the given string. + +**Parameters**: + +- `name`: The name of the rule to remove. + +**Returns**: `Some(ContentRule)` if found and removed, otherwise `None`. + +**Example**: + +```rust +let removed = config.remove_rule_by_name("blog"); +``` + +#### `has_duplicate_rule_names` + +```rust +#[must_use] +pub fn has_duplicate_rule_names(&self) -> bool +``` + +**Purpose**: Checks whether any two content rules share the same `name`. + +**Returns**: `true` if duplicates exist, `false` otherwise. + +**Implementation**: Uses a `HashSet` to detect duplicates; O(n) time. + +**Example**: + +```rust +if config.has_duplicate_rule_names() { + // handle error +} +``` + +### Config::validate + +```rust +pub fn validate(&self) -> Result<()> +``` + +**Purpose**: Performs comprehensive validation of the configuration. Returns `Ok(())` if the configuration is valid, otherwise an `Err(Error::Validation(...))` with a descriptive message. + +**Validation Rules**: + +1. `site.site_name` must not be empty or whitespace‑only. +2. At least one content rule must be defined. +3. No duplicate content rule names. +4. For each content rule (indexed from 0): + - `name` must not be empty or whitespace‑only. + - `pattern` must not be empty or whitespace‑only. + - `pattern` must not contain the substring `".."` (to prevent path traversal). + - `template` must not be empty or whitespace‑only. +5. If `site.base_url` is `Some`, it must start with `"http://"` or `"https://"`. + +**Returns**: + +- `Ok(())` if all checks pass. +- `Err(Error::Validation(message))` on the first failure encountered. + +**Example** (from tests): + +```rust +let config = valid_config(); // has one rule +assert!(config.validate().is_ok()); +``` + +### Config Serialization + +The `Config` struct can be serialized to and deserialized from YAML and JSON via convenience methods. + +#### `from_yaml_str` + +```rust +pub fn from_yaml_str(yaml: &str) -> Result +``` + +**Purpose**: Parses a YAML string into a `Config`. + +**Parameters**: + +- `yaml`: YAML content as a string. + +**Returns**: + +- `Ok(Config)` on success. +- `Err(Error::Config)` if the YAML is invalid (with the underlying `serde_yaml` error message included). + +**Example**: + +```rust +let config = Config::from_yaml_str("site:\n site_name: Test\n")?; +``` + +#### `to_yaml_string` + +```rust +pub fn to_yaml_string(&self) -> Result +``` + +**Purpose**: Serializes the `Config` to a YAML string. + +**Returns**: + +- `Ok(String)` with YAML representation. +- `Err(Error::Serialization)` if serialization fails. + +#### `from_json_str` + +```rust +pub fn from_json_str(json: &str) -> Result +``` + +**Purpose**: Parses a JSON string into a `Config`. + +**Parameters**: + +- `json`: JSON content as a string. + +**Returns**: + +- `Ok(Config)` on success. +- `Err(Error::Config)` if the JSON is invalid. + +#### `to_json_string` + +```rust +pub fn to_json_string(&self) -> Result +``` + +**Purpose**: Serializes the `Config` to a JSON string. + +**Returns**: + +- `Ok(String)` with JSON representation. +- `Err(Error::Serialization)` on failure. + +**Example roundtrip**: + +```rust +let yaml = config.to_yaml_string()?; +let parsed = Config::from_yaml_str(&yaml)?; +assert_eq!(config, parsed); +``` + +--- + +## Error Handling + +The `Config` methods and `validate` use `librawssg_error::Result` (alias for `std::result::Result`). The relevant error variants used in this crate are: + +- `Error::Config` – For YAML/JSON deserialization failures. +- `Error::Serialization` – For serialization failures. +- `Error::Validation` – For `validate()` failures. + +All error messages are descriptive and include context (e.g., which rule is invalid, what condition was violated). + +--- + +## Serialization Details + +All configuration structs derive `Serialize` and `Deserialize` from `serde`. Default values are applied during deserialization for missing fields via `#[serde(default = "function")]` or `#[serde(default)]` (which uses `Default::default()` for the field type). + +- `BuildConfig`: Each field has a custom default function. +- `ContentRule`: Optional fields use `#[serde(default)]`. +- `NavItem`: No special defaults; all fields are required in input, but `Default` is derived for programmatic creation. +- `SiteConfig`: `site_name` has a custom default, `language` has a custom default returning `Some("en")`, others use `#[serde(default)]`. +- `Config`: `site` and `build` are required in YAML/JSON (they don't have `#[serde(default)]` at the field level, but the struct itself derives `Default` and the fields are not marked optional; however, when deserializing a top‑level `Config`, missing `site` or `build` will cause an error because they are not optional. In practice, configuration files should include these sections or rely on the `Default` implementation when constructing programmatically). **Important**: The `Config` struct's fields are not marked with `#[serde(default)]`, so during deserialization, missing `site` or `build` will cause a parse error. Users must provide at least `site` and `build` keys (can be empty maps to get defaults via the inner structs' own defaults). The `content_rules` and `extra` fields have `#[serde(default)]` so they can be omitted. + +--- + +## Examples from Tests + +The test suite provides extensive examples for each type. Below are selected snippets. + +### BuildConfig + +```rust +let build = BuildConfig::default(); +assert_eq!(build.output_dir, "dist"); + +let yaml = "content_dir: custom_content\noutput_dir: public\n"; +let build: BuildConfig = serde_yaml::from_str(yaml)?; +assert_eq!(build.content_dir, "custom_content"); +assert_eq!(build.templates_dir, "templates"); +``` + +### ContentRule + +```rust +let mut rule = ContentRule::new("page", "**/*.html", "base"); +rule.list_enabled = true; +rule.list_template = Some("list".into()); +rule.extra.insert("key".into(), json!("value")); +``` + +### NavItem + +```rust +let mut parent = NavItem::new("Docs", "/docs"); +parent.children.push(NavItem::new("API", "/docs/api")); +``` + +### SiteConfig + +```rust +let site = SiteConfig::new("My Site"); +assert_eq!(site.language.as_deref(), Some("en")); +``` + +### Config Validation + +```rust +let mut config = Config::new().with_site_name("My Site"); +config.add_content_rule(ContentRule::new("page", "**/*.html", "base")); +assert!(config.validate().is_ok()); + +config.site.base_url = Some("ftp://example.com".to_string()); +assert!(config.validate().is_err()); +``` + +--- + +## Testing Suite Overview + +The crate includes five test files: + +- `build_tests.rs` – Tests `BuildConfig` defaults, `new()`, and YAML roundtrip. +- `config_tests.rs` – Extensive tests for `Config`: creation, rule management, validation (all rules), YAML/JSON roundtrip, and error cases. +- `content_rule_tests.rs` – Tests `ContentRule` constructor, defaults, and serialization. +- `nav_tests.rs` – Tests `NavItem` constructor, defaults, children, and serialization. +- `site_tests.rs` – Tests `SiteConfig` constructor, defaults, and serialization. + +All tests are self‑contained and use the `must!` macro to unwrap results with a helpful message on failure. They serve as executable examples of the API usage. + +--- + +## Conclusion + +`librawssg_config` provides a clean and extensible configuration system for a static site generator. With sensible defaults, comprehensive validation, and full serde support, it covers the needs of both simple and complex site configurations. The types are designed for ergonomic use and can be easily loaded from YAML or JSON files, making it straightforward to define site‑wide settings, build paths, navigation, and content processing rules. + +For further details, refer to the source code and test files. diff --git a/librawssg_demo/README.md b/librawssg_demo/README.md new file mode 100644 index 0000000..75b8ae0 --- /dev/null +++ b/librawssg_demo/README.md @@ -0,0 +1,259 @@ +# librawssg_demo + +A complete demo application for **librawssg**, a modular static site generator framework written in Rust. This project showcases how to assemble the various `librawssg` crates into a working static site generator that reads raw HTML fragments, renders them through a Tera template, copies static assets, and outputs a fully static website. + +--- + +## Table of Contents + +1. [Overview](#overview) +2. [Features](#features) +3. [Project Structure](#project-structure) +4. [Prerequisites](#prerequisites) +5. [Installation & Build](#installation--build) +6. [Running the Demo](#running-the-demo) +7. [How It Works](#how-it-works) + - [The `RawFileProcessor`](#the-rawfileprocessor) + - [Template Rendering with Tera](#template-rendering-with-tera) + - [Configuration](#configuration) + - [Static Assets](#static-assets) + - [Output Generation](#output-generation) +8. [Customization Guide](#customization-guide) + - [Adding New Content Files](#adding-new-content-files) + - [Changing the Template](#changing-the-template) + - [Adding Custom Processors](#adding-custom-processors) + - [Adding Generators](#adding-generators) +9. [Underlying Crates](#underlying-crates) +10. [Troubleshooting](#troubleshooting) +11. [License](#license) + +--- + +## Overview + +`librawssg_demo` demonstrates a minimal but functional static site generator built with the `librawssg` framework. It uses: + +- **`librawssg_config`** for configuration management. +- **`librawssg_handler`** to define a custom `Processor` that handles `.raw` files. +- **`librawssg_templates`** for Tera-based rendering. +- **`librawssg_fs`** for filesystem abstraction (using `RealFs`). +- **`librawssg_compiler`** to orchestrate the build pipeline. + +The demo processes `.raw` files (containing HTML fragments) from `src/content`, renders them using a Tera template (`base.tera`), copies static files from `src/static`, and outputs the final site into the `dist/` folder. + +--- + +## Features + +- **Modular architecture**: Each component (processing, rendering, file I/O, configuration) is separated and replaceable. +- **Custom content processing**: The included `RawFileProcessor` reads `.raw` files and converts them into `Document` objects. +- **Template rendering**: Uses the Tera template engine with a `TeraContextBuilder` to inject page data. +- **Static asset copying**: Automatically copies the `static` directory to the output. +- **Atomic output**: The build process writes to a temporary directory and atomically replaces the final output. +- **Extensible**: Easily add new processors, renderers, context builders, or generators. + +--- + +## Project Structure + +``` +librawssg_demo/ +├── Cargo.toml +└── src/ + ├── content/ + │ ├── about.raw + │ └── index.raw + ├── static/ + │ └── style.css + ├── templates/ + │ └── base.tera + └── main.rs +``` + +- **`Cargo.toml`** – Defines the package and dependencies on the `librawssg_*` crates (via path). +- **`src/main.rs`** – Entry point; configures and runs the pipeline. +- **`src/content/`** – Contains source content files (`.raw`). These are processed by `RawFileProcessor`. +- **`src/static/`** – Static assets (CSS, images, etc.) that are copied verbatim to the output. +- **`src/templates/`** – Tera templates used for rendering. +- **`dist/`** – Generated output (created at runtime, not stored in version control). + +--- + +## Prerequisites + +- **Rust toolchain** (stable, edition 2024) – Install via [rustup](https://rustup.rs/). +- **Cargo** – Comes with Rust. +- The `librawssg_*` crates must be available at the relative paths specified in `Cargo.toml`. This demo assumes a workspace layout where the crates are siblings of `librawssg_demo`. + +--- + +## Installation & Build + +1. **Clone the repository** (or navigate to the demo directory inside the workspace). + +2. **Build the project**: + + ```bash + cargo build + ``` + +3. **Run the demo**: + ```bash + cargo run + ``` + +Upon successful execution, the output site will be generated in the `dist/` folder (relative to the project root). + +--- + +## Running the Demo + +Execute: + +```bash +cargo run +``` + +The program will: + +1. Read the configuration (created programmatically in `main.rs`). +2. Load templates from `src/templates`. +3. Process all `.raw` files in `src/content`. +4. Render each document using the `base.tera` template. +5. Copy static files from `src/static`. +6. Write everything into `dist/` atomically. + +After completion, you can open `dist/index.html` in a browser to view the generated site. + +--- + +## How It Works + +### The `RawFileProcessor` + +The `RawFileProcessor` implements the `Processor` trait from `librawssg_handler`. Its responsibilities: + +- **`can_process`**: Returns `true` only for files with a `.raw` extension. +- **`process`**: + - Reads the file content using the provided `FileSystem`. + - Derives a title from the file stem (e.g., `my-page` becomes `My Page`). + - Creates a `Metadata` object with the title. + - Constructs a `Document` with: + - `body`: the raw HTML content (will be inserted into the template). + - `url`: the same relative path but with `.html` extension. + - `output_path`: the relative output path (same as URL). + - `content_type`: hardcoded to `"page"`. + +### Template Rendering with Tera + +- **Renderer**: `TeraRenderer` (from `librawssg_templates`) is used. It loads all templates from the `src/templates` directory recursively. +- **Context Builder**: `TeraContextBuilder` creates a `tera::Context` containing: + - `site`: the full `SiteConfig`. + - `page_title`, `page_description`, `page_author`, etc. + - `page_content`: the raw body of the document (inserted with `| safe` filter in the template). +- **Template**: `base.tera` defines the overall HTML structure. It uses `{{ page_title }}` and `{{ page_content | safe }}` to inject data. + +### Configuration + +In `main.rs`, a `Config` object is built: + +- `site_name` is set to `"Demo Site"`. +- One `ContentRule` is added: `name = "page"`, `pattern = "**/*.raw"`, `template = "base.tera"`. +- Build paths (`content_dir`, `output_dir`, `static_dir`) are set to absolute paths under the project root. + +### Static Assets + +The `config.build.static_dir` points to `src/static`. During generation, the pipeline copies this directory recursively to the output. In this demo, the `style.css` file will appear at `dist/static/style.css`. The template references it via `/style.css` assuming the static directory is copied with its name preserved (i.e., `static/`). To make the CSS load correctly, the static directory is copied as `dist/static/`, and the link in the template should be `/static/style.css`. However, the provided template uses `/style.css` – this might be a slight mismatch. In a real deployment you may want to adjust either the static directory name or the link. The demo as provided will copy `static/` into `dist/static/`, but the HTML link expects `/style.css` which would 404 unless the server serves the static directory at root. For local file viewing, it won't work directly. This is a common issue; you can either change the static dir to be copied to the root (by setting `static_dir` to a path with no subdirectory) or modify the template link to `/static/style.css`. The demo code keeps the default behavior; user may need to adjust. + +### Output Generation + +The `Pipeline` orchestrates the build: + +1. Processes all content files. +2. Groups documents by content type. +3. Renders non-list documents. +4. (No list pages in this demo because `list_enabled` is not set.) +5. Copies static assets. +6. Runs any custom generators (none in this demo). +7. Atomically replaces the `dist/` directory. + +--- + +## Customization Guide + +### Adding New Content Files + +Simply create a new `.raw` file in `src/content/`. For example, `src/content/contact.raw`: + +```html +

Contact

+

Email us at hello@example.com

+``` + +The next `cargo run` will automatically process it and generate `dist/contact.html`. + +### Changing the Template + +Edit `src/templates/base.tera`. You can use any Tera syntax. The available context variables include: + +- `page_title` +- `page_content` (raw HTML) +- `page_description` +- `page_author` +- `page_date` +- `page_tags` +- `page_url` +- `page_depth` +- `page_type` +- `page_is_list` +- `page_list_items` +- `site` (the full `SiteConfig`) + +You can also add custom fields to `Metadata` or use the `extra` maps. + +### Adding Custom Processors + +Implement the `Processor` trait and add it to the pipeline builder with `.add_processor(Box::new(MyProcessor))`. For example, you could create a Markdown processor that handles `.md` files. + +### Adding Generators + +Implement the `Generator` trait and add it with `.add_generator(Box::new(MyGenerator))`. Generators run after documents and static files are written, allowing you to create RSS feeds, sitemaps, or search indexes. + +--- + +## Underlying Crates + +This demo relies on the following `librawssg` crates (all located in sibling directories): + +- [`librawssg_compiler`](../librawssg_compiler) – Pipeline and builder. +- [`librawssg_config`](../librawssg_config) – Configuration types. +- [`librawssg_fs`](../librawssg_fs) – Filesystem abstraction. +- [`librawssg_handler`](../librawssg_handler) – `Document`, `Metadata`, and `Processor` trait. +- [`librawssg_templates`](../librawssg_templates) – Rendering traits and Tera implementation. +- [`librawssg_error`](../librawssg_error) – Unified error types. + +All are used via path dependencies, so they must be present in the workspace. + +--- + +## Troubleshooting + +- **Build errors about missing crates** + Ensure all `librawssg_*` crates are checked out at the correct relative paths (siblings to this demo). Check `Cargo.toml` for the `path` attributes. + +- **Static CSS not loading** + By default, the static directory is copied with its name (`static/`). The template currently links to `/style.css`. Either: + - Change the template link to `/static/style.css`, or + - Modify `config.build.static_dir` to a path whose basename is empty (e.g., copy contents directly to output root) by adjusting the pipeline's static copying logic (not recommended for this demo). + +- **`cargo run` fails with permission errors** + The output directory `dist/` may already exist and be locked. Try deleting it manually or ensuring you have write permissions. + +- **Tera template syntax errors** + Validate your template syntax. The error messages from Tera are descriptive and will indicate the problem line. + +--- + +## License + +This demo is licensed under the **MIT License**. See the `LICENSE` file in the repository root for details. diff --git a/librawssg_demo/src/static/style.css b/librawssg_demo/src/static/style.css index 458e2aa..1a9201e 100644 --- a/librawssg_demo/src/static/style.css +++ b/librawssg_demo/src/static/style.css @@ -1,9 +1,9 @@ body { - font-family: sans-serif; - max-width: 800px; - margin: 0 auto; - padding: 2rem; + font-family: sans-serif; + max-width: 800px; + margin: 0 auto; + padding: 2rem; } nav a { - margin-right: 1rem; + margin-right: 1rem; } diff --git a/librawssg_error/README.md b/librawssg_error/README.md index 07aee98..76e814b 100644 --- a/librawssg_error/README.md +++ b/librawssg_error/README.md @@ -1,4 +1,685 @@ # librawssg_error -Unified error types for the librawssg static site generator ecosystem. -Provides a single `Error` enum and `Result` alias used by all other crates. +**Version**: 1.0.0 (implied) +**Crate name**: `librawssg_error` +**Description**: Defines a comprehensive error enum and `Result` alias for use across the `librawssg` static site generator ecosystem. The error type is built with `thiserror` for ergonomic `Display` and `Error` implementations, supports source chaining, and is designed to cover all common failure modes in SSG operations. + +--- + +## Table of Contents + +1. [Overview](#overview) +2. [Dependencies](#dependencies) +3. [The `Error` Enum](#the-error-enum) + - [Enum Definition](#enum-definition) + - [Attributes and Derives](#attributes-and-derives) + - [Non‑Exhaustive](#non-exhaustive) +4. [Variants](#variants) + - [`Io`](#io) + - [`Config`](#config) + - [`Metadata`](#metadata) + - [`Render`](#render) + - [`Processor`](#processor) + - [`Generator`](#generator) + - [`PathTraversal`](#pathtraversal) + - [`MissingConfig`](#missingconfig) + - [`Generation`](#generation) + - [`NotFound`](#notfound) + - [`Serialization`](#serialization) + - [`Validation`](#validation) + - [`Duplicate`](#duplicate) + - [`InvalidState`](#invalidstate) + - [`Internal`](#internal) +5. [`Result` Type Alias](#resultt-type-alias) +6. [Error Sources and `std::error::Error`](#error-sources-and-stderroerror) +7. [Conversion from `std::io::Error`](#conversion-from-stdioerror) +8. [Usage Examples from Tests](#usage-examples-from-tests) + - [Display Messages](#display-messages) + - [Using the `?` Operator](#using-the--operator) + - [Source Chain](#source-chain) + - [Property Tests](#property-tests) +9. [Guidelines for Error Usage](#guidelines-for-error-usage) +10. [Testing Suite Overview](#testing-suite-overview) +11. [Conclusion](#conclusion) + +--- + +## Overview + +The `librawssg_error` crate provides a single, unified error type for the entire static site generator ecosystem. Instead of having each module define its own error types, they all share this `Error` enum, which categorizes failures into well‑defined variants. The enum is derived with `thiserror::Error`, giving each variant an automatic `Display` implementation based on a custom message pattern, and an automatic `std::error::Error` implementation that preserves source chains when applicable. + +The crate also exports a `Result` type alias, simplifying function signatures throughout the codebase. + +Key features: + +- **Rich error categories** – 15 distinct variants covering I/O, configuration, parsing, rendering, processing, generation, security, and internal errors. +- **Source chaining** – The `Metadata` variant can wrap an underlying error (e.g., a YAML parsing error) and expose it via `source()`. +- **Convenient conversion** – `From` allows using the `?` operator directly in functions returning `Result`. +- **Non‑exhaustive** – The enum is marked `#[non_exhaustive]`, enabling future additions without breaking downstream code. + +--- + +## Dependencies + +- `thiserror` – Provides the `#[derive(Error)]` macro that generates `Display` and `Error` implementations from the attributes. +- `core::error::Error` (or `std::error::Error`) – Used as a trait object for the `source` field in the `Metadata` variant. + +No other external crates are required. + +--- + +## The `Error` Enum + +### Enum Definition + +```rust +use core::error::Error as CoreError; +use std::path::PathBuf; +use thiserror::Error; + +#[derive(Debug, Error)] +#[non_exhaustive] +pub enum Error { + #[error("I/O error: {0}")] + Io(#[from] std::io::Error), + + #[error("Configuration error: {0}")] + Config(String), + + #[error("Failed to parse metadata in {path}")] + Metadata { + path: PathBuf, + #[source] + source: Box, + }, + + #[error("Template rendering error: {0}")] + Render(String), + + #[error("Content processor error: {0}")] + Processor(String), + + #[error("Generator error: {0}")] + Generator(String), + + #[error("Path traversal attempt detected: {0}")] + PathTraversal(String), + + #[error("Missing configuration key: {0}")] + MissingConfig(String), + + #[error("Site generation error: {0}")] + Generation(String), + + #[error("Resource not found: {0}")] + NotFound(String), + + #[error("Serialization error: {0}")] + Serialization(String), + + #[error("Validation error: {0}")] + Validation(String), + + #[error("Duplicate value: {0}")] + Duplicate(String), + + #[error("Invalid state: {0}")] + InvalidState(String), + + #[error("Internal error: {0}")] + Internal(String), +} +``` + +### Attributes and Derives + +- **`#[derive(Debug, Error)]`** – Derives `Debug` and `std::error::Error`. The `Error` derive from `thiserror` also generates a `Display` implementation based on the `#[error("...")]` attributes. +- **`#[non_exhaustive]`** – Indicates that the enum may gain new variants in future releases. Downstream crates must not exhaustively match on this enum; they must include a wildcard arm (`_`) when matching. + +### Non‑Exhaustive + +Because the enum is non‑exhaustive, external code cannot write: + +```rust +match err { + Error::Io(_) => ..., + Error::Config(_) => ..., + // all variants... +} +``` + +without including a catch‑all arm: + +```rust +match err { + Error::Io(_) => ..., + Error::Config(_) => ..., + // ... + _ => { /* handle unknown future variants */ } +} +``` + +This ensures forward compatibility. + +--- + +## Variants + +### `Io` + +```rust +#[error("I/O error: {0}")] +Io(#[from] std::io::Error), +``` + +- **Description**: Wraps a standard library I/O error. Used for any filesystem operation failure (reading, writing, deleting, etc.). +- **Fields**: Contains a single `std::io::Error`. +- **Display**: `"I/O error: {underlying_io_error_message}"`. +- **`#[from]`**: Automatically provides `From for Error`, allowing the `?` operator in functions returning `Result`. +- **Source**: `source()` returns `Some(&io_error)` because `std::io::Error` implements `std::error::Error`. + +**Example**: + +```rust +let io_err = std::io::Error::new(std::io::ErrorKind::NotFound, "file missing"); +let err = Error::Io(io_err); +assert_eq!(err.to_string(), "I/O error: file missing"); +``` + +--- + +### `Config` + +```rust +#[error("Configuration error: {0}")] +Config(String), +``` + +- **Description**: Indicates a problem with configuration data (e.g., invalid YAML, malformed settings). +- **Fields**: A `String` containing a human‑readable description. +- **Display**: `"Configuration error: {message}"`. +- **Source**: `None` (no underlying error is stored). + +**Example**: + +```rust +let err = Error::Config("invalid YAML".to_string()); +assert_eq!(err.to_string(), "Configuration error: invalid YAML"); +``` + +--- + +### `Metadata` + +```rust +#[error("Failed to parse metadata in {path}")] +Metadata { + path: PathBuf, + #[source] + source: Box, +}, +``` + +- **Description**: Used when parsing metadata (e.g., front matter) fails. Stores the path of the problematic file and the original error. +- **Fields**: + - `path: PathBuf` – The path to the source file where metadata parsing failed. + - `source: Box` – The underlying error that caused the failure (boxed trait object). +- **Display**: `"Failed to parse metadata in {path}"`. The `{path}` placeholder prints the `PathBuf` using its `Display` implementation. +- **Source**: `source()` returns `Some(&*source)` (the boxed error as a `&dyn Error`), enabling error chain inspection. +- **Note**: The `#[source]` attribute tells `thiserror` to use this field as the error source. The `Box` allows storing any error type that is `Send + Sync`. + +**Example**: + +```rust +use std::path::PathBuf; + +let source: Box = + Box::new(std::io::Error::other("bad yaml")); +let err = Error::Metadata { + path: PathBuf::from("content/post.md"), + source, +}; + +assert_eq!(err.to_string(), "Failed to parse metadata in content/post.md"); +let as_core_error: &dyn core::error::Error = &err; +let source_ref = as_core_error.source().unwrap(); +assert_eq!(source_ref.to_string(), "bad yaml"); +``` + +--- + +### `Render` + +```rust +#[error("Template rendering error: {0}")] +Render(String), +``` + +- **Description**: Signifies an error during template rendering (e.g., missing variable, template not found, syntax error). +- **Fields**: A `String` describing the rendering problem. +- **Display**: `"Template rendering error: {message}"`. +- **Source**: `None`. + +**Example**: + +```rust +let err = Error::Render("template not found".to_string()); +assert_eq!(err.to_string(), "Template rendering error: template not found"); +``` + +--- + +### `Processor` + +```rust +#[error("Content processor error: {0}")] +Processor(String), +``` + +- **Description**: Used when a content processor (e.g., Markdown parser, Sass compiler) fails. +- **Fields**: A `String` with details about the processor failure. +- **Display**: `"Content processor error: {message}"`. +- **Source**: `None`. + +**Example**: + +```rust +let err = Error::Processor("custom processor failed".to_string()); +assert_eq!(err.to_string(), "Content processor error: custom processor failed"); +``` + +--- + +### `Generator` + +```rust +#[error("Generator error: {0}")] +Generator(String), +``` + +- **Description**: Represents an error in a generator component (e.g., RSS feed generation, sitemap creation). +- **Fields**: A `String` describing the generator error. +- **Display**: `"Generator error: {message}"`. +- **Source**: `None`. + +**Example**: + +```rust +let err = Error::Generator("RSS generation failed".to_string()); +assert_eq!(err.to_string(), "Generator error: RSS generation failed"); +``` + +--- + +### `PathTraversal` + +```rust +#[error("Path traversal attempt detected: {0}")] +PathTraversal(String), +``` + +- **Description**: Indicates a path traversal attack was attempted or a path escapes a safe root directory. +- **Fields**: A `String` containing the offending path or description. +- **Display**: `"Path traversal attempt detected: {message}"`. +- **Source**: `None`. + +**Example**: + +```rust +let err = Error::PathTraversal("../escape".to_string()); +assert_eq!(err.to_string(), "Path traversal attempt detected: ../escape"); +``` + +--- + +### `MissingConfig` + +```rust +#[error("Missing configuration key: {0}")] +MissingConfig(String), +``` + +- **Description**: Signals that a required configuration key is absent. +- **Fields**: A `String` naming the missing key. +- **Display**: `"Missing configuration key: {key}"`. +- **Source**: `None`. + +**Example**: + +```rust +let err = Error::MissingConfig("base_url".to_string()); +assert_eq!(err.to_string(), "Missing configuration key: base_url"); +``` + +--- + +### `Generation` + +```rust +#[error("Site generation error: {0}")] +Generation(String), +``` + +- **Description**: A general error during the site generation phase (e.g., failed to write output). +- **Fields**: A `String` with more information. +- **Display**: `"Site generation error: {message}"`. +- **Source**: `None`. + +**Example**: + +```rust +let err = Error::Generation("output write failed".to_string()); +assert_eq!(err.to_string(), "Site generation error: output write failed"); +``` + +--- + +### `NotFound` + +```rust +#[error("Resource not found: {0}")] +NotFound(String), +``` + +- **Description**: Used when a requested resource (file, asset, page) cannot be found. +- **Fields**: A `String` identifying the missing resource. +- **Display**: `"Resource not found: {resource}"`. +- **Source**: `None`. + +**Example**: + +```rust +let err = Error::NotFound("asset.css".to_string()); +assert_eq!(err.to_string(), "Resource not found: asset.css"); +``` + +--- + +### `Serialization` + +```rust +#[error("Serialization error: {0}")] +Serialization(String), +``` + +- **Description**: Indicates a failure during serialization or deserialization (e.g., JSON conversion error). +- **Fields**: A `String` describing the serialization problem. +- **Display**: `"Serialization error: {message}"`. +- **Source**: `None`. + +**Example**: + +```rust +let err = Error::Serialization("invalid JSON".to_string()); +assert_eq!(err.to_string(), "Serialization error: invalid JSON"); +``` + +--- + +### `Validation` + +```rust +#[error("Validation error: {0}")] +Validation(String), +``` + +- **Description**: Represents a validation failure (e.g., invalid input, constraint violation). +- **Fields**: A `String` explaining what failed validation. +- **Display**: `"Validation error: {message}"`. +- **Source**: `None`. + +**Example**: + +```rust +let err = Error::Validation("name too long".to_string()); +assert_eq!(err.to_string(), "Validation error: name too long"); +``` + +--- + +### `Duplicate` + +```rust +#[error("Duplicate value: {0}")] +Duplicate(String), +``` + +- **Description**: Signals that a duplicate value was encountered where uniqueness was expected (e.g., duplicate key in a map). +- **Fields**: A `String` identifying the duplicated item. +- **Display**: `"Duplicate value: {message}"`. +- **Source**: `None`. + +**Example**: + +```rust +let err = Error::Duplicate("duplicate key".to_string()); +assert_eq!(err.to_string(), "Duplicate value: duplicate key"); +``` + +--- + +### `InvalidState` + +```rust +#[error("Invalid state: {0}")] +InvalidState(String), +``` + +- **Description**: Indicates an unexpected program state (e.g., a null where a value is required, inconsistent internal data). +- **Fields**: A `String` describing the invalid state. +- **Display**: `"Invalid state: {message}"`. +- **Source**: `None`. + +**Example**: + +```rust +let err = Error::InvalidState("unexpected null".to_string()); +assert_eq!(err.to_string(), "Invalid state: unexpected null"); +``` + +--- + +### `Internal` + +```rust +#[error("Internal error: {0}")] +Internal(String), +``` + +- **Description**: Used for internal errors that should not normally occur (e.g., bugs in the code, unreachable conditions). +- **Fields**: A `String` with details suitable for debugging. +- **Display**: `"Internal error: {message}"`. +- **Source**: `None`. + +**Example**: + +```rust +let err = Error::Internal("bug in code".to_string()); +assert_eq!(err.to_string(), "Internal error: bug in code"); +``` + +--- + +## `Result` Type Alias + +```rust +pub type Result = core::result::Result; +``` + +- **Purpose**: A convenient alias so that functions can return `Result` instead of the more verbose `std::result::Result`. +- **Usage**: Throughout the `librawssg` ecosystem, functions that may fail with any of the above errors use this alias. + +**Example**: + +```rust +fn read_config(path: &str) -> librawssg_error::Result { + let content = std::fs::read_to_string(path)?; // `?` converts io::Error into Error::Io + Ok(content) +} +``` + +--- + +## Error Sources and `std::error::Error` + +All variants of `Error` implement `std::error::Error` (via `thiserror`). The `source()` method returns: + +- For `Io`: `Some(&self.0)` (the underlying `io::Error`). +- For `Metadata`: `Some(self.source.as_ref())` (the boxed error). +- For all other variants: `None`. + +This allows error chains to be inspected using `std::error::Error::source()`. + +**Example** (from integration tests): + +```rust +let source: Box = + Box::new(std::io::Error::other("bad yaml")); +let err = Error::Metadata { + path: PathBuf::from("content/post.md"), + source, +}; + +let as_core_error: &dyn core::error::Error = &err; +let source_ref = as_core_error.source().unwrap(); +assert_eq!(source_ref.to_string(), "bad yaml"); +``` + +--- + +## Conversion from `std::io::Error` + +The `Io` variant has the `#[from]` attribute, which automatically generates: + +```rust +impl From for Error { + fn from(err: std::io::Error) -> Error { + Error::Io(err) + } +} +``` + +This enables the `?` operator to convert `std::io::Error` into `Error` in any function returning `Result` (or `librawssg_error::Result`). + +**Example**: + +```rust +fn read_file(path: &str) -> Result { + let content = std::fs::read_to_string(path)?; // io::Error becomes Error::Io + Ok(content) +} +``` + +--- + +## Usage Examples from Tests + +The test suite provides excellent examples of how to construct and use the error type. + +### Display Messages + +Each variant has a test asserting its exact `Display` output. For example: + +```rust +#[test] +fn display_for_config_error() { + let err = Error::Config("invalid YAML".to_string()); + assert_eq!(err.to_string(), "Configuration error: invalid YAML"); +} +``` + +All 15 variants have similar tests in `unit_tests.rs`. + +### Using the `?` Operator + +The integration test `io_error_propagates_via_question_mark` demonstrates how `?` works: + +```rust +fn read_file(path: &str) -> Result { + let content = std::fs::read_to_string(path)?; + Ok(content) +} + +#[test] +fn io_error_propagates_via_question_mark() { + let result = read_file("definitely_not_exists.txt"); + assert!(result.is_err()); + assert!(matches!(result, Err(Error::Io(_)))); +} +``` + +### Source Chain + +The `metadata_error_can_hold_boxed_dyn_error` test shows how to store an arbitrary error and retrieve it via `source()`: + +```rust +let source: Box = + Box::new(std::io::Error::other("bad yaml")); +let err = Error::Metadata { + path: PathBuf::from("content/post.md"), + source, +}; + +let as_core_error: &dyn core::error::Error = &err; +let source_ref = as_core_error.source(); +assert!(source_ref.is_some()); +``` + +### Property Tests + +Property tests verify that the error message always preserves the input string exactly, regardless of content (including empty strings, newlines, special characters): + +```rust +#[test] +fn config_error_message_preserves_input() { + let samples = [ + "", + "short", + "a very long error message with symbols !@#$%^&*()", + "line1\nline2", + ]; + + for sample in samples { + let err = Error::Config(sample.to_string()); + assert_eq!(err.to_string(), format!("Configuration error: {sample}")); + } +} +``` + +Similar tests exist for `PathTraversal` and `Render`. + +--- + +## Guidelines for Error Usage + +When writing code in the `librawssg` ecosystem, follow these recommendations: + +- **Use the most specific variant** that describes the failure. For example: + - I/O failures → `Error::Io`. + - Missing file/resource → `Error::NotFound`. + - Invalid user input → `Error::Validation`. + - Security issue (path traversal) → `Error::PathTraversal`. +- **Attach context when possible** – Include the relevant path, key, or identifier in the error message string. +- **Preserve source errors** – If an underlying error is available, use the `Metadata` variant (or add a new variant with a `#[source]` field) to maintain the error chain. +- **Avoid matching exhaustively on `Error`** – Because the enum is non‑exhaustive, always include a catch‑all arm when matching to prevent future breakage. +- **Use `Result` alias** for concise function signatures. + +--- + +## Testing Suite Overview + +The crate includes three test files: + +- **`unit_tests.rs`** – Tests each variant’s `Display` message, `source()` for `Metadata`, `From` conversion, the `Result` alias, and `Debug` output. +- **`integration_tests.rs`** – Tests the `?` operator integration and the source chain for `Metadata` using a boxed dynamic error. +- **`property_tests.rs`** – Property‑based tests that verify error messages preserve arbitrary input strings for `Config`, `PathTraversal`, and `Render`. + +Together, these tests ensure the error type is robust, easy to use, and consistent. + +--- + +## Conclusion + +`librawssg_error` provides a centralized, well‑structured error type for the entire static site generator project. With 15 descriptive variants, automatic `Display` and `Error` implementations, convenient conversion from `io::Error`, and support for error sources, it simplifies error handling across all modules. The non‑exhaustive design guarantees future extensibility without breaking downstream code. + +For further details, refer to the source code and test files. diff --git a/librawssg_fs/README.md b/librawssg_fs/README.md index 3ca0a9b..e5f42de 100644 --- a/librawssg_fs/README.md +++ b/librawssg_fs/README.md @@ -1,4 +1,867 @@ # librawssg_fs -Filesystem abstraction layer for librawssg. -Defines the `FileSystem` trait and provides a real implementation. +**Version**: 1.0.0 (implied) +**Crate name**: `librawssg_fs` +**Description**: A filesystem abstraction layer for static site generators. Defines the `FileSystem` trait with a comprehensive set of file and directory operations, and provides a concrete implementation `RealFs` that delegates to `std::fs` and `walkdir`. The trait includes built‑in path traversal protection and convenience methods for atomic operations. + +--- + +## Table of Contents + +1. [Overview](#overview) +2. [Modules](#modules) +3. [Trait `FileSystem`](#trait-filesystem) + - [Trait Definition](#trait-definition) + - [Required Methods](#required-methods) + - [`read_to_string`](#read_to_string) + - [`read_bytes`](#read_bytes) + - [`write`](#write) + - [`create_dir_all`](#create_dir_all) + - [`remove_dir_all`](#remove_dir_all) + - [`remove_file`](#remove_file) + - [`create_dir`](#create_dir) + - [`exists`](#exists) + - [`is_dir`](#is_dir) + - [`is_file`](#is_file) + - [`read_dir`](#read_dir) + - [`copy_file`](#copy_file) + - [`copy_dir_all`](#copy_dir_all) + - [`walk_dir`](#walk_dir) + - [`canonicalize`](#canonicalize) + - [`rename`](#rename) + - [`atomic_write`](#atomic_write) + - [`touch`](#touch) + - [`metadata`](#metadata) + - [`symlink_metadata`](#symlink_metadata) + - [`permissions`](#permissions) + - [`set_permissions`](#set_permissions) + - [`read_link`](#read_link) + - [`hard_link`](#hard_link) + - [Provided (Default) Methods](#provided-default-methods) + - [`is_symlink`](#is_symlink) + - [`canonicalize_or_join`](#canonicalize_or_join) + - [`safe_join`](#safe_join) + - [`copy`](#copy) + - [`rename_or_copy`](#rename_or_copy) +4. [Struct `RealFs`](#struct-realfs) + - [Implementation Details](#implementation-details) + - [Example Usage](#example-usage) +5. [Error Handling](#error-handling) +6. [Implementing a Custom `FileSystem`](#implementing-a-custom-filesystem) +7. [Security Considerations](#security-considerations) +8. [Testing Suite Overview](#testing-suite-overview) +9. [Complete Code Examples from Tests](#complete-code-examples-from-tests) + +--- + +## Overview + +`librawssg_fs` provides a trait‑based abstraction over filesystem operations. This allows static site generator components to interact with the filesystem without being tightly coupled to `std::fs`. It enables: + +- **Testability**: Mock filesystems can be injected in unit tests. +- **Portability**: Different filesystem backends (e.g., in‑memory, virtual) can implement the trait. +- **Security**: Built‑in path traversal protection through `safe_join` and `canonicalize_or_join`. + +The crate exports: + +- `pub trait FileSystem` – The main abstraction. +- `pub struct RealFs` – A zero‑sized type that implements `FileSystem` using the real OS filesystem. + +--- + +## Modules + +The crate root (`lib.rs`) defines the `FileSystem` trait and re‑exports `RealFs` from the `real` module. + +```rust +pub mod real; +pub use real::RealFs; +``` + +There is also an internal module `real.rs` containing the `RealFs` implementation. + +--- + +## Trait `FileSystem` + +The `FileSystem` trait is the core of this crate. It is object‑safe and requires implementors to be `Send + Sync` (safe to share across threads). The trait provides many required methods and several methods with default implementations. + +```rust +pub trait FileSystem: Send + Sync { + // Required methods (see below) + // Provided methods with default implementations +} +``` + +### Required Methods + +These methods **must** be implemented by any type that implements `FileSystem`. They map closely to `std::fs` functions and `walkdir` functionality. + +#### `read_to_string` + +```rust +fn read_to_string(&self, path: &Path) -> io::Result; +``` + +**Purpose**: Reads the entire contents of a file into a `String`. + +**Parameters**: + +- `path`: The path to the file to read. + +**Returns**: `Ok(String)` containing the file contents, or an `Err(io::Error)` if the file cannot be read (e.g., not found, permission denied, invalid UTF‑8). + +**Example**: + +```rust +let content = fs.read_to_string(Path::new("hello.txt"))?; +``` + +#### `read_bytes` + +```rust +fn read_bytes(&self, path: &Path) -> io::Result>; +``` + +**Purpose**: Reads the entire contents of a file as raw bytes. + +**Parameters**: + +- `path`: The path to the file. + +**Returns**: `Ok(Vec)` with the file bytes, or an `Err(io::Error)`. + +**Example**: + +```rust +let data = fs.read_bytes(Path::new("image.png"))?; +``` + +#### `write` + +```rust +fn write(&self, path: &Path, content: &[u8]) -> io::Result<()>; +``` + +**Purpose**: Writes the given bytes to a file, creating any necessary parent directories. + +**Parameters**: + +- `path`: Destination file path. +- `content`: Bytes to write. + +**Returns**: `Ok(())` on success, or `Err(io::Error)` on failure (e.g., permission denied, disk full). + +**Behavior**: The default `RealFs` implementation creates parent directories before writing (via `create_dir_all` on the parent). This is convenient for writing deeply nested outputs. + +**Example**: + +```rust +fs.write(Path::new("a/b/c.txt"), b"hello")?; +``` + +#### `create_dir_all` + +```rust +fn create_dir_all(&self, path: &Path) -> io::Result<()>; +``` + +**Purpose**: Creates a directory and all its missing parents. + +**Parameters**: + +- `path`: The directory path to create. + +**Returns**: `Ok(())` or `Err(io::Error)`. + +**Note**: Unlike `create_dir`, this does **not** error if the directory already exists. + +**Example**: + +```rust +fs.create_dir_all(Path::new("a/b/c"))?; +``` + +#### `remove_dir_all` + +```rust +fn remove_dir_all(&self, path: &Path) -> io::Result<()>; +``` + +**Purpose**: Removes a directory and all its contents recursively. + +**Parameters**: + +- `path`: Directory path to remove. + +**Returns**: `Ok(())` or `Err(io::Error)` (e.g., directory does not exist, permission denied). + +**Warning**: This is destructive and cannot be undone. + +#### `remove_file` + +```rust +fn remove_file(&self, path: &Path) -> io::Result<()>; +``` + +**Purpose**: Deletes a single file. + +**Parameters**: + +- `path`: File path to remove. + +**Returns**: `Ok(())` or `Err(io::Error)`. + +#### `create_dir` + +```rust +fn create_dir(&self, path: &Path) -> io::Result<()>; +``` + +**Purpose**: Creates a single directory. Fails if the parent directory does not exist or if the directory already exists. + +**Parameters**: + +- `path`: Directory path to create. + +**Returns**: `Ok(())` or `Err(io::Error)` (e.g., already exists, parent missing). + +#### `exists` + +```rust +fn exists(&self, path: &Path) -> bool; +``` + +**Purpose**: Checks whether a path exists (as a file, directory, symlink, etc.). + +**Parameters**: + +- `path`: Path to check. + +**Returns**: `true` if the path exists, `false` otherwise. + +**Note**: This method does not follow symlinks for broken symlinks; it returns `false` for a broken symlink. + +#### `is_dir` + +```rust +fn is_dir(&self, path: &Path) -> bool; +``` + +**Purpose**: Checks whether the path points to a directory. + +**Returns**: `true` if it is a directory, `false` otherwise (including if it does not exist). + +#### `is_file` + +```rust +fn is_file(&self, path: &Path) -> bool; +``` + +**Purpose**: Checks whether the path points to a regular file. + +**Returns**: `true` if it is a regular file, `false` otherwise. + +#### `read_dir` + +```rust +fn read_dir(&self, path: &Path) -> io::Result>; +``` + +**Purpose**: Lists all entries (files and directories) directly inside a directory. + +**Parameters**: + +- `path`: Directory path. + +**Returns**: `Ok(Vec)` containing the full paths of all entries, or `Err(io::Error)`. + +**Note**: The order is not guaranteed. It does not recurse into subdirectories. + +#### `copy_file` + +```rust +fn copy_file(&self, from: &Path, to: &Path) -> io::Result; +``` + +**Purpose**: Copies a file from `from` to `to`. If `to` already exists, it will be overwritten. + +**Parameters**: + +- `from`: Source file path. +- `to`: Destination file path. + +**Returns**: `Ok(u64)` with the number of bytes copied, or `Err(io::Error)`. + +**Note**: Does not create parent directories of `to` in the default `RealFs`; use `copy` or `copy_dir_all` for that. + +#### `copy_dir_all` + +```rust +fn copy_dir_all(&self, from: &Path, to: &Path) -> io::Result<()>; +``` + +**Purpose**: Recursively copies a directory tree from `from` to `to`. Creates the destination directory and all parent directories as needed. + +**Parameters**: + +- `from`: Source directory path. +- `to`: Destination directory path. + +**Returns**: `Ok(())` or `Err(io::Error)`. + +**Behavior**: + +1. Creates `to` directory. +2. Walks all files in `from` (using `walk_dir`). +3. For each file, computes relative path and creates parent directories in `to`, then copies the file. + +#### `walk_dir` + +```rust +fn walk_dir(&self, root: &Path) -> io::Result>; +``` + +**Purpose**: Recursively collects all **files** under `root`. Does not include directories or symlinks to directories. + +**Parameters**: + +- `root`: Root directory to traverse. + +**Returns**: `Ok(Vec)` with the full paths of all files, or `Err(io::Error)`. + +**Note**: The default `RealFs` uses the `walkdir` crate to handle traversal. It follows symlinks? (The `WalkDir::new` default does not follow symlinks; it will include symlinks but not traverse into them unless `.follow_links(true)` is set. Here symlinks to files will be included? `entry.file_type().is_file()` will be true for a symlink to a file? Actually `file_type()` returns the type of the symlink itself, not the target, unless `follow_links` is used. So symlinks are not considered files and are skipped.) + +#### `canonicalize` + +```rust +fn canonicalize(&self, path: &Path) -> io::Result; +``` + +**Purpose**: Returns the canonical, absolute form of a path, resolving all symbolic links and normalizing `.` and `..` components. + +**Parameters**: + +- `path`: The path to canonicalize. + +**Returns**: `Ok(PathBuf)` with the canonical path, or `Err(io::Error)` (e.g., path does not exist). + +#### `rename` + +```rust +fn rename(&self, from: &Path, to: &Path) -> io::Result<()>; +``` + +**Purpose**: Renames (moves) a file or directory from `from` to `to`. On most filesystems this is an atomic operation when source and destination are on the same filesystem. + +**Parameters**: + +- `from`: Source path. +- `to`: Destination path. + +**Returns**: `Ok(())` or `Err(io::Error)`. + +**Note**: If `to` exists, it may be overwritten (platform‑dependent). Does not work across different mount points (returns `CrossesDevices` error). + +#### `atomic_write` + +```rust +fn atomic_write(&self, path: &Path, content: &[u8]) -> io::Result<()>; +``` + +**Purpose**: Writes data to a file atomically by first writing to a temporary file and then renaming it over the target path. + +**Parameters**: + +- `path`: Destination file path. +- `content`: Bytes to write. + +**Returns**: `Ok(())` or `Err(io::Error)`. + +**Behavior**: + +1. Creates a temporary file with extension `.tmp` (by calling `with_extension("tmp")` on the target path). +2. Writes the content to the temporary file (using `write`, which creates parent directories). +3. Renames the temporary file to the target path (using `rename`). +4. If the rename fails, attempts to remove the temporary file and returns the error. + +**Note**: The temporary file name is derived from the target; it is not a hidden file and may collide if multiple writes happen concurrently to the same path. This is a best‑effort atomic write suitable for many use cases. + +#### `touch` + +```rust +fn touch(&self, path: &Path) -> io::Result<()>; +``` + +**Purpose**: Creates an empty file at `path` or updates its access/modification timestamp if it already exists. + +**Parameters**: + +- `path`: File path. + +**Returns**: `Ok(())` or `Err(io::Error)`. + +**Behavior**: + +1. Creates parent directories (like `write`). +2. Opens the file in append/create mode, which creates it if missing. +3. Calls `sync_all()` to flush to disk (optional, but ensures metadata is updated). + +**Note**: Existing file content is preserved. + +#### `metadata` + +```rust +fn metadata(&self, path: &Path) -> io::Result; +``` + +**Purpose**: Returns metadata for a file or directory, following symlinks. + +**Parameters**: + +- `path`: Path to query. + +**Returns**: `Ok(fs::Metadata)` or `Err(io::Error)`. + +#### `symlink_metadata` + +```rust +fn symlink_metadata(&self, path: &Path) -> io::Result; +``` + +**Purpose**: Returns metadata for a path **without** following symlinks (i.e., metadata of the symlink itself). + +**Parameters**: + +- `path`: Path to query. + +**Returns**: `Ok(fs::Metadata)` or `Err(io::Error)`. + +#### `permissions` + +```rust +fn permissions(&self, path: &Path) -> io::Result; +``` + +**Purpose**: Reads the permissions of a file or directory. + +**Parameters**: + +- `path`: Path to query. + +**Returns**: `Ok(fs::Permissions)` or `Err(io::Error)`. + +**Note**: The default `RealFs` obtains permissions from `metadata`, which follows symlinks. + +#### `set_permissions` + +```rust +fn set_permissions(&self, path: &Path, permissions: std::fs::Permissions) -> io::Result<()>; +``` + +**Purpose**: Sets the permissions of a file or directory. + +**Parameters**: + +- `path`: Target path. +- `permissions`: New permissions. + +**Returns**: `Ok(())` or `Err(io::Error)`. + +#### `read_link` + +```rust +fn read_link(&self, path: &Path) -> io::Result; +``` + +**Purpose**: Reads the target of a symbolic link. + +**Parameters**: + +- `path`: Path to the symlink. + +**Returns**: `Ok(PathBuf)` containing the link target, or `Err(io::Error)` if the path is not a symlink or does not exist. + +#### `hard_link` + +```rust +fn hard_link(&self, from: &Path, to: &Path) -> io::Result<()>; +``` + +**Purpose**: Creates a hard link from `from` to `to`. + +**Parameters**: + +- `from`: Existing file path. +- `to`: New hard link path. + +**Returns**: `Ok(())` or `Err(io::Error)`. + +**Note**: Both paths must be on the same filesystem. + +--- + +### Provided (Default) Methods + +These methods have default implementations that rely on the required methods. Implementors may override them for performance or platform‑specific behavior. + +#### `is_symlink` + +```rust +fn is_symlink(&self, path: &Path) -> bool { + self.symlink_metadata(path) + .is_ok_and(|meta| meta.file_type().is_symlink()) +} +``` + +**Purpose**: Checks whether the given path is a symbolic link. + +**Returns**: `true` if the path is a symlink (even if broken), `false` otherwise. + +**Implementation**: Uses `symlink_metadata` (which does not follow symlinks) and checks the file type. + +**Example**: + +```rust +if fs.is_symlink(Path::new("link")) { ... } +``` + +#### `canonicalize_or_join` + +```rust +fn canonicalize_or_join(&self, base: &Path, candidate: &Path) -> io::Result +``` + +**Purpose**: Safely resolves a possibly non‑existent path relative to `base`. If the joined path exists, it is canonicalized; otherwise it returns the canonical parent joined with the file name, after normalizing `.` and `..` components. + +**Parameters**: + +- `base`: The base directory (usually already canonical). +- `candidate`: A relative path (may contain `.` and `..`). + +**Returns**: `Ok(PathBuf)` with the resolved path, or `Err(io::Error)` if path traversal is detected or other errors occur. + +**Detailed Behavior**: + +1. Normalizes the `candidate` path by iterating over its components: + - `CurDir` (`.`) is ignored. + - `ParentDir` (`..`) causes the last normal component to be popped. If there is no previous normal component (i.e., attempt to go above root), it returns `PermissionDenied` with message `"path traversal detected"`. + - `Prefix` and `RootDir` components are pushed (though they are unusual for relative candidates and may cause issues later). + - `Normal` components are pushed. +2. Joins the normalized candidate with `base`. +3. If the joined path exists, canonicalizes it (resolving symlinks, etc.). +4. If it does not exist: + - Canonicalizes the parent directory of the joined path. + - Appends the file name of the joined path to the canonical parent. + - Returns that path. + +**Security**: This method prevents `..` from escaping the base directory (unless there are symlinks that point outside; canonicalization of existing paths can still lead outside base, which is why `safe_join` adds an extra check). For non‑existent paths, the parent canonicalization ensures that the final path is within the canonical base. + +**Example** (from tests): + +```rust +let existing = base.join("existing.txt"); +fs.write(&existing, b"data")?; + +let canon_existing = fs.canonicalize_or_join(base, Path::new("existing.txt"))?; +let canon_direct = fs.canonicalize(&existing)?; +assert_eq!(canon_existing, canon_direct); + +let missing = Path::new("missing.txt"); +let canon_missing = fs.canonicalize_or_join(base, missing)?; +let canon_base = fs.canonicalize(base)?; +assert_eq!(canon_missing, canon_base.join(missing)); +``` + +#### `safe_join` + +```rust +fn safe_join(&self, base: &Path, candidate: &Path) -> io::Result +``` + +**Purpose**: Safely joins a candidate path to a base directory, ensuring the result is **within** the base directory (no path traversal). This is the recommended way to compute destination paths for user‑provided or untrusted relative paths. + +**Parameters**: + +- `base`: The base directory (can be relative; it will be canonicalized internally). +- `candidate`: A relative path (may contain `.` and `..`). + +**Returns**: `Ok(PathBuf)` with the resolved path, guaranteed to start with the canonical base. Returns `Err(io::Error)` with `PermissionDenied` if the resolved path escapes the base (e.g., `candidate = "../secret"`). + +**Implementation Details**: + +1. Canonicalizes `base`. +2. Calls `canonicalize_or_join` with the canonical base and `candidate`. +3. Checks that the resulting path starts with the canonical base. If not, returns `PermissionDenied`. + +**Why needed**: Although `canonicalize_or_join` prevents simple `..` traversal, symlinks inside the base directory could cause a resolved path to point outside the base even after normalization. The `starts_with` check enforces containment. + +**Example**: + +```rust +let base = tmp.path().join("base"); +fs.create_dir_all(&base)?; + +let safe = fs.safe_join(&base, Path::new("inside.txt"))?; +assert!(safe.starts_with(&base)); + +let traversal = Path::new("../escape.txt"); +let result = fs.safe_join(&base, traversal); +assert!(result.is_err()); +``` + +#### `copy` + +```rust +fn copy(&self, from: &Path, to: &Path) -> io::Result<()> +``` + +**Purpose**: Copies a file or directory from `from` to `to`. If `from` is a directory, it recursively copies the whole tree; if it is a file, it performs a single file copy. + +**Parameters**: + +- `from`: Source path. +- `to`: Destination path. + +**Returns**: `Ok(())` or `Err(io::Error)`. + +**Implementation**: + +```rust +if self.is_dir(from) { + self.copy_dir_all(from, to) +} else { + self.copy_file(from, to).map(|_| ()) +} +``` + +**Note**: Does not create parent directories of `to` for file copies (unless `copy_file` implementation does; the default `RealFs::copy_file` does not). For directories, `copy_dir_all` does create `to` and parents as needed. + +**Example**: + +```rust +fs.copy(&src_file, &dst_file)?; +fs.copy(&src_dir, &dst_dir)?; +``` + +#### `rename_or_copy` + +```rust +fn rename_or_copy(&self, from: &Path, to: &Path) -> io::Result<()> +``` + +**Purpose**: Attempts to rename `from` to `to`. If the rename fails with `ErrorKind::CrossesDevices` (i.e., source and destination are on different filesystems), it falls back to copying the directory tree and then removing the source. + +**Parameters**: + +- `from`: Source path. +- `to`: Destination path. + +**Returns**: `Ok(())` or `Err(io::Error)`. + +**Behavior**: + +1. Try `rename(from, to)`. +2. If success, return `Ok(())`. +3. If error kind is `CrossesDevices`: + - `copy_dir_all(from, to)` to copy contents. + - `remove_dir_all(from)` to delete source. + - Return `Ok(())`. +4. Otherwise, return the original error. + +**Note**: The fallback only works for directories (as the code uses `copy_dir_all` and `remove_dir_all`). For a file across devices, this will likely fail. This method is useful for moving directories across mount points. + +**Example**: + +```rust +fs.rename_or_copy(&src_dir, &dst_dir)?; +``` + +--- + +## Struct `RealFs` + +`RealFs` is a zero‑sized struct that implements `FileSystem` by delegating directly to the operating system’s filesystem APIs. + +```rust +#[derive(Debug, Default, Clone, Copy)] +pub struct RealFs; +``` + +It has no fields and can be instantiated with `RealFs` or `RealFs::default()`. + +### Implementation Details + +`RealFs` uses: + +- `std::fs` for most operations. +- `walkdir::WalkDir` for `walk_dir`. +- The `tracing::instrument` attribute is applied to most methods for logging (though `tracing` is not enabled by default; it can be used with a subscriber). + +All methods follow the behavior described in the trait definitions. The `write` method creates parent directories before writing, and `atomic_write` uses a temporary `.tmp` file. + +### Example Usage + +```rust +use librawssg_fs::{FileSystem, RealFs}; +use std::path::Path; + +let fs = RealFs; + +// Write a file +fs.write(Path::new("output/file.txt"), b"Hello")?; + +// Read it back +let content = fs.read_to_string(Path::new("output/file.txt"))?; +assert_eq!(content, "Hello"); + +// Create directory +fs.create_dir_all(Path::new("output/sub"))?; + +// Copy directory +fs.copy_dir_all(Path::new("output"), Path::new("backup"))?; +``` + +--- + +## Error Handling + +All methods that can fail return `io::Result` (i.e., `Result`). This is the standard error type from the standard library, so no custom error enum is defined in this crate. Consumers can inspect the error kind (e.g., `ErrorKind::NotFound`, `PermissionDenied`, `CrossesDevices`) to handle specific failures. + +The provided security methods (`canonicalize_or_join` and `safe_join`) return `io::Error` with `ErrorKind::PermissionDenied` when path traversal is detected, along with the message `"path traversal detected"`. + +--- + +## Implementing a Custom `FileSystem` + +To create a mock filesystem or an alternative backend, implement the `FileSystem` trait. You must provide implementations for all 24 required methods. The provided methods can be left as default unless you need custom behavior. + +**Example of a minimal mock** (from tests, adapted): + +```rust +use librawssg_fs::FileSystem; +use std::io; +use std::path::{Path, PathBuf}; + +struct DummyFs; + +impl FileSystem for DummyFs { + fn read_to_string(&self, _path: &Path) -> io::Result { + Err(io::Error::other("not implemented")) + } + // ... implement all other required methods similarly + // (returning Err or trivial values) +} +``` + +Because `FileSystem` is `Send + Sync`, your mock must also be thread‑safe. In practice, you can use `Arc` or interior mutability if state is needed. + +--- + +## Security Considerations + +The library includes two methods specifically designed to prevent path traversal attacks: + +- **`canonicalize_or_join`**: Normalizes `..` and `.` and ensures the final path does not go above the base (in terms of lexical components). However, it may still follow symlinks that point outside the base if the path exists. +- **`safe_join`**: Combines `canonicalize_or_join` with a `starts_with` check on the canonical base, providing a stronger guarantee that the result is contained within the base directory. + +**Recommendation**: Always use `safe_join` when constructing output paths from untrusted input (e.g., user‑supplied relative URLs). Avoid using `join` directly followed by canonicalization without containment checks. + +--- + +## Testing Suite Overview + +The test file `tests/filesystem.rs` contains comprehensive tests for `RealFs` and the provided methods. It uses `tempfile::TempDir` to create isolated temporary directories. The tests cover: + +- Basic read/write operations (string and bytes) +- Creating directories and files +- Error cases for non‑existent paths +- `copy_file`, `copy_dir_all`, `rename`, `remove_file`, `remove_dir_all` +- `atomic_write` (including nested paths and overwriting) +- `touch` (creating new and preserving existing content) +- `walk_dir` (recursive collection) +- `canonicalize_or_join` (existing and missing paths) +- `safe_join` (rejecting traversal, allowing dot segments inside) +- Metadata and permissions +- Symlink and hard link operations (Unix only) + +All tests can be run with `cargo test`. + +--- + +## Complete Code Examples from Tests + +Below are selected examples from the test suite that illustrate common usage patterns. They can be copied and adapted. + +### Writing and Reading a String + +```rust +use librawssg_fs::{FileSystem, RealFs}; +use std::path::Path; +use tempfile::TempDir; + +let tmp = TempDir::new().unwrap(); +let fs = RealFs; +let file_path = tmp.path().join("hello.txt"); + +fs.write(&file_path, b"world").unwrap(); +let content = fs.read_to_string(&file_path).unwrap(); +assert_eq!(content, "world"); +``` + +### Atomic Write Overwriting + +```rust +let file = tmp.path().join("atomic.txt"); +fs.atomic_write(&file, b"first").unwrap(); +fs.atomic_write(&file, b"second").unwrap(); +let content = fs.read_to_string(&file).unwrap(); +assert_eq!(content, "second"); +``` + +### Safe Join Blocking Traversal + +```rust +let base = tmp.path().join("base"); +fs.create_dir_all(&base).unwrap(); + +let safe = fs.safe_join(&base, Path::new("inside.txt")).unwrap(); +assert!(safe.starts_with(&base)); + +let traversal = Path::new("../escape.txt"); +assert!(fs.safe_join(&base, traversal).is_err()); +``` + +### Copying a Directory Recursively + +```rust +let src = tmp.path().join("src_dir"); +let dst = tmp.path().join("dst_dir"); +fs.create_dir_all(&src.join("nested")).unwrap(); +fs.write(&src.join("file1.txt"), b"one").unwrap(); +fs.write(&src.join("nested").join("file2.txt"), b"two").unwrap(); + +fs.copy_dir_all(&src, &dst).unwrap(); +assert!(fs.exists(&dst.join("file1.txt"))); +assert!(fs.exists(&dst.join("nested").join("file2.txt"))); +``` + +### Using `walk_dir` to Gather All Files + +```rust +let root = tmp.path().join("root"); +fs.create_dir_all(&root.join("sub")).unwrap(); +fs.write(&root.join("root.txt"), b"root").unwrap(); +fs.write(&root.join("sub").join("sub.txt"), b"sub").unwrap(); + +let files = fs.walk_dir(&root).unwrap(); +assert_eq!(files.len(), 2); +``` + +--- + +## Summary + +`librawssg_fs` provides a robust, thread‑safe filesystem abstraction with built‑in path traversal protection and convenience methods for atomic operations and cross‑device moves. The `RealFs` implementation is ready to use, and the trait enables easy mocking for unit tests. The extensive test suite validates all features and serves as living documentation. + +For any additional details, refer to the source code and inline comments. diff --git a/librawssg_handler/README.md b/librawssg_handler/README.md index 581174e..f61dde5 100644 --- a/librawssg_handler/README.md +++ b/librawssg_handler/README.md @@ -1,4 +1,775 @@ # librawssg_handler -Content processing contracts for librawssg. -Defines the `Processor` trait and related types for turning source files into documents. +**Version**: 1.0.0 (implied) +**Crate name**: `librawssg_handler` +**Description**: Core data structures and traits for building a static site generator (SSG) handler. Provides `Document`, `Metadata`, and a `Processor` trait for processing content files. + +--- + +## Table of Contents + +1. [Overview](#overview) +2. [Modules](#modules) +3. [Struct `Document`](#struct-document) + - [Fields](#document-fields) + - [Constructor `new()`](#document-new) + - [Method `relative_url()`](#document-relative_url) + - [Method `add_taxonomy()`](#document-add_taxonomy) + - [Method `depth()`](#document-depth) + - [Method `with_list_items()`](#document-with_list_items) +4. [Struct `Metadata`](#struct-metadata) + - [Fields](#metadata-fields) + - [Constructor `new()`](#metadata-new) + - [Method `is_draft()`](#metadata-is_draft) + - [Method `insert_extra()`](#metadata-insert_extra) + - [Method `get_extra()`](#metadata-get_extra) + - [Serialization & Deserialization](#metadata-serialization) + - [Default Implementation](#metadata-default) +5. [Trait `Processor`](#trait-processor) + - [Required Methods](#processor-required-methods) + - [Provided Methods](#processor-provided-methods) + - [Implementing the Trait](#processor-implementation) +6. [Error Handling](#error-handling) +7. [External Traits & Types](#external-traits-and-types) +8. [Examples from Tests](#examples-from-tests) +9. [Validation Rules Summary](#validation-rules-summary) +10. [Testing Suite Overview](#testing-suite-overview) + +--- + +## Overview + +`librawssg_handler` is the core library for a static site generator. It defines the essential data structures used to represent a processed document (`Document`) and its front matter (`Metadata`). Additionally, it provides a pluggable `Processor` trait that allows different file types to be processed into `Document` instances. + +The crate is intended to be used in conjunction with: + +- `librawssg_error`: Provides the `Error` and `Result` types for consistent error handling. +- `librawssg_fs`: Defines a `FileSystem` trait abstracting file I/O operations (used by `Processor`). + +--- + +## Modules + +The library is organized into three public modules: + +- **`document`** – Contains the `Document` struct. +- **`metadata`** – Contains the `Metadata` struct. +- **`processor`** – Contains the `Processor` trait. + +All public types are re‑exported at the crate root for convenience: + +```rust +pub use document::Document; +pub use metadata::Metadata; +pub use processor::Processor; +``` + +--- + +## Struct `Document` + +Represents a fully processed content item ready for rendering or further processing. + +```rust +#[derive(Debug, Clone, PartialEq, Serialize)] +#[non_exhaustive] +pub struct Document { + pub metadata: Metadata, + pub body: String, + pub url: String, + pub output_path: PathBuf, + pub source_path: PathBuf, + pub depth: usize, + pub content_type: String, + pub is_list: bool, + pub list_items: Option>, + pub taxonomies: HashMap>, +} +``` + +### Document Fields + +| Field | Type | Description | +| -------------- | ------------------------------ | -------------------------------------------------------------------------------------------------- | +| `metadata` | `Metadata` | Front matter metadata associated with the document. | +| `body` | `String` | The processed content body (e.g., rendered HTML, Markdown text, etc.). | +| `url` | `String` | The relative URL where the document will be accessible (e.g., `"blog/my-post.html"`). | +| `output_path` | `PathBuf` | Filesystem path where the final output file should be written (e.g., `"blog/my-post/index.html"`). | +| `source_path` | `PathBuf` | Path to the original source file (e.g., `"content/blog/my-post.md"`). | +| `depth` | `usize` | Depth of the document in the site hierarchy (0 for top‑level). Used for sorting or navigation. | +| `content_type` | `String` | Identifier for the kind of content (e.g., `"blog"`, `"page"`, `"article"`). | +| `is_list` | `bool` | Indicates whether this document represents a list of other documents (e.g., an index page). | +| `list_items` | `Option>` | If `is_list` is true, may contain the child documents. `None` otherwise or when not set. | +| `taxonomies` | `HashMap>` | A map of taxonomy names (e.g., `"categories"`, `"tags"`) to lists of terms. | + +> **Note:** The `#[non_exhaustive]` attribute means that external crates cannot exhaustively match on `Document` or construct it with a struct literal; they must use the provided constructor or update syntax. This allows adding fields in the future without breaking downstream code. + +### Document::new + +```rust +pub fn new( + metadata: Metadata, + body: impl Into, + url: impl Into, + output_path: impl Into, + source_path: impl Into, + depth: usize, + content_type: impl Into, + is_list: bool, +) -> Result +``` + +**Purpose**: Creates a new `Document` after validating the provided arguments. + +**Parameters**: + +- `metadata`: A fully constructed `Metadata` instance. +- `body`: The content body (accepts any type convertible to `String`). +- `url`: The desired relative URL (must not be empty or whitespace only). +- `output_path`: The target output path (must not be empty). +- `source_path`: The source file path (must have a file name component). +- `depth`: The hierarchy depth (must be ≤ 1000). +- `content_type`: A string identifying the content type (e.g., `"blog"`, `"page"`). +- `is_list`: Boolean indicating whether this document is a list container. + +**Returns**: + +- `Ok(Document)` on success. +- `Err(librawssg_error::Error::Validation(message))` if any validation rule fails. + +**Validation Rules**: + +1. `url` must not be empty or contain only whitespace. +2. `output_path` must not be empty (as an OS string). +3. `source_path` must have a file name (i.e., its last component is not `..` or empty). +4. `depth` must not exceed 1000. + +**Example**: + +```rust +use librawssg_handler::{Document, Metadata}; +use std::path::PathBuf; + +let metadata = Metadata::new("My Post", "A short description")?; + +let document = Document::new( + metadata, + "

Hello

World

", + "blog/my-post.html", + "blog/my-post/index.html", + "content/blog/my-post.md", + 1, + "blog", + false, +)?; +``` + +--- + +### Document::relative_url + +```rust +#[must_use] +pub fn relative_url(&self) -> &str +``` + +**Purpose**: Returns the relative URL of the document. + +**Returns**: A string slice referencing the `url` field. + +**Example**: + +```rust +let doc = /* ... */; +assert_eq!(doc.relative_url(), "blog/my-post.html"); +``` + +--- + +### Document::add_taxonomy + +```rust +pub fn add_taxonomy(&mut self, name: impl Into, items: Vec) +``` + +**Purpose**: Inserts or replaces a taxonomy entry in the document’s `taxonomies` map. + +**Parameters**: + +- `name`: Taxonomy name (converted into `String`). +- `items`: A vector of string terms belonging to that taxonomy. + +**Behavior**: If a taxonomy with the same name already exists, its value is replaced. + +**Example**: + +```rust +let mut doc = /* ... */; +doc.add_taxonomy("categories", vec!["rust".to_string(), "ssg".to_string()]); +assert_eq!(doc.taxonomies["categories"], vec!["rust", "ssg"]); +``` + +--- + +### Document::depth + +```rust +#[must_use] +pub const fn depth(&self) -> usize +``` + +**Purpose**: Returns the `depth` field. + +**Returns**: The document’s depth as a `usize`. + +**Example**: + +```rust +let doc = /* ... */; +assert_eq!(doc.depth(), 1); +``` + +--- + +### Document::with_list_items + +```rust +#[must_use] +pub fn with_list_items(mut self, items: Vec) -> Self +``` + +**Purpose**: Consumes the document, sets its `list_items` field to `Some(items)`, and returns the modified document. + +**Parameters**: + +- `items`: A vector of `Document` instances that are children of this list document. + +**Returns**: The same document with `list_items` set. + +**Example**: + +```rust +let parent = Document::new(/* ... */)?; +let child1 = Document::new(/* ... */)?; +let child2 = Document::new(/* ... */)?; +let list_doc = parent.with_list_items(vec![child1, child2]); +assert!(list_doc.list_items.is_some()); +``` + +--- + +## Struct `Metadata` + +Represents front matter (metadata) for a document. The struct is serializable and deserializable, making it suitable for parsing from formats like YAML or TOML front matter. + +```rust +#[derive(Debug, Clone, PartialEq, Eq, Serialize, Deserialize, Default)] +#[non_exhaustive] +pub struct Metadata { + pub title: String, + pub description: String, + pub author: Option, + pub repo_url: Option, + pub license: Option, + pub date: Option, + pub updated: Option, + pub tags: Vec, + pub draft: bool, + pub extra: HashMap, +} +``` + +### Metadata Fields + +| Field | Type | Description | +| ------------- | ------------------------------------ | ---------------------------------------------------------------------------- | +| `title` | `String` | The document title (required, cannot be empty). | +| `description` | `String` | A short description of the content. | +| `author` | `Option` | The author’s name, if known. | +| `repo_url` | `Option` | URL to the source repository. | +| `license` | `Option` | License identifier (e.g., `"MIT"`, `"Apache-2.0"`). | +| `date` | `Option` | Publication date (ISO 8601 date, e.g., `2026-09-08`). | +| `updated` | `Option` | Last modification date. | +| `tags` | `Vec` | List of tags (keywords) associated with the content. | +| `draft` | `bool` | If true, the document is considered a draft and may be excluded from builds. | +| `extra` | `HashMap` | Arbitrary extra key–value pairs for custom metadata. | + +> **Note:** `#[non_exhaustive]` prevents exhaustive struct literals outside the crate; use the provided constructors or update syntax. + +### Metadata::new + +```rust +pub fn new( + title: impl Into, + description: impl Into, +) -> librawssg_error::Result +``` + +**Purpose**: Creates a `Metadata` instance with a required `title` and `description`. All other fields are set to their default values. + +**Parameters**: + +- `title`: The title (must not be empty or whitespace only). +- `description`: A description string. + +**Returns**: + +- `Ok(Metadata)` on success. +- `Err(librawssg_error::Error::Validation("metadata title cannot be empty"))` if the title is empty or whitespace. + +**Example**: + +```rust +use librawssg_handler::Metadata; + +let meta = Metadata::new("My Title", "My Description")?; +assert_eq!(meta.title, "My Title"); +assert!(!meta.draft); +assert!(meta.tags.is_empty()); +``` + +--- + +### Metadata::is_draft + +```rust +#[must_use] +pub const fn is_draft(&self) -> bool +``` + +**Purpose**: Returns the `draft` field. + +**Returns**: `true` if the document is marked as a draft, otherwise `false`. + +**Example**: + +```rust +let mut meta = Metadata::new("Title", "Desc")?; +assert!(!meta.is_draft()); +meta.draft = true; +assert!(meta.is_draft()); +``` + +--- + +### Metadata::insert_extra + +```rust +pub fn insert_extra(&mut self, key: impl Into, value: impl Into) +``` + +**Purpose**: Inserts or updates an entry in the `extra` map. + +**Parameters**: + +- `key`: The key (converted to `String`). +- `value`: Any type convertible to `serde_json::Value` (e.g., strings, numbers, booleans, arrays, objects, or `serde_json::json!` macro results). + +**Behavior**: If the key already exists, its value is overwritten. + +**Example**: + +```rust +use serde_json::json; + +let mut meta = Metadata::new("Title", "Desc")?; +meta.insert_extra("key", "value"); +meta.insert_extra("number", 42); +meta.insert_extra("flag", true); +meta.insert_extra("nested", json!({"foo": "bar"})); +``` + +--- + +### Metadata::get_extra + +```rust +#[must_use] +pub fn get_extra(&self, key: &str) -> Option<&serde_json::Value> +``` + +**Purpose**: Retrieves a reference to the value stored under `key` in the `extra` map. + +**Parameters**: + +- `key`: The key to look up. + +**Returns**: + +- `Some(&Value)` if the key exists. +- `None` otherwise. + +**Example**: + +```rust +let meta = /* ... */; +if let Some(v) = meta.get_extra("key") { + assert_eq!(v, &json!("value")); +} +``` + +--- + +### Metadata Serialization + +`Metadata` derives both `Serialize` and `Deserialize`, so it can be converted to/from JSON, YAML, etc. This is particularly useful for reading front matter from source files. + +**Serialization Example**: + +```rust +use serde_json; + +let meta = Metadata::new("Hello", "World")?; +let json_str = serde_json::to_string(&meta)?; +// {"title":"Hello","description":"World","author":null,...} +``` + +**Deserialization Example**: + +```rust +let json_str = r#"{ + "title": "Hello", + "description": "World", + "author": "Alice", + "date": "2026-09-08", + "tags": ["rust", "ssg"], + "draft": false, + "extra": {"foo": "bar"} +}"#; + +let meta: Metadata = serde_json::from_str(json_str)?; +assert_eq!(meta.author.as_deref(), Some("Alice")); +``` + +--- + +### Metadata Default + +The `Default` trait is implemented. All fields are set to sensible empty values: + +- `title`: empty string +- `description`: empty string +- `author`, `repo_url`, `license`, `date`, `updated`: `None` +- `tags`: empty vector +- `draft`: `false` +- `extra`: empty `HashMap` + +**Example**: + +```rust +let meta = Metadata::default(); +assert_eq!(meta.title, ""); +assert!(!meta.draft); +assert!(meta.tags.is_empty()); +``` + +--- + +## Trait `Processor` + +The `Processor` trait defines an interface for components that can transform a source file into a `Document`. Multiple processors may be registered and invoked based on their ability to handle a given file. + +```rust +pub trait Processor: Send + Sync { + fn name(&self) -> &str; + + fn priority(&self) -> i32 { + 0 + } + + fn can_process(&self, relative_path: &Path, original_path: &Path) -> bool; + + fn process( + &self, + fs: &dyn FileSystem, + relative_path: &Path, + content_dir: &Path, + ) -> Result>; +} +``` + +### Processor Required Methods + +#### `name()` + +```rust +fn name(&self) -> &str +``` + +**Purpose**: Returns a human‑readable identifier for the processor (e.g., `"markdown"`, `"sass"`). + +#### `can_process()` + +```rust +fn can_process(&self, relative_path: &Path, original_path: &Path) -> bool +``` + +**Purpose**: Determines whether this processor should handle the given file. + +**Parameters**: + +- `relative_path`: The path of the file relative to the content directory. +- `original_path`: The full original path (often the same as `content_dir.join(relative_path)`). + +**Returns**: `true` if the processor can process this file; `false` otherwise. + +**Typical Implementation**: Check file extension or other attributes. + +**Example**: + +```rust +fn can_process(&self, relative_path: &Path, _original_path: &Path) -> bool { + relative_path.extension().and_then(|e| e.to_str()) == Some("md") +} +``` + +#### `process()` + +```rust +fn process( + &self, + fs: &dyn FileSystem, + relative_path: &Path, + content_dir: &Path, +) -> Result> +``` + +**Purpose**: Reads the source file, processes it, and returns an optional `Document`. + +**Parameters**: + +- `fs`: A reference to a `FileSystem` implementation for performing I/O operations. +- `relative_path`: Path of the source file relative to the content directory. +- `content_dir`: The root directory containing all source content. + +**Returns**: + +- `Ok(Some(document))` if processing succeeded and produced a document. +- `Ok(None)` if the processor decides not to produce a document (e.g., the file is ignored). +- `Err(librawssg_error::Error)` if an error occurred during processing. + +**Note**: The `FileSystem` trait is defined in the `librawssg_fs` crate. It abstracts many common file operations, allowing processors to be tested with mock filesystems. + +--- + +### Processor Provided Methods + +#### `priority()` + +```rust +fn priority(&self) -> i32 { + 0 +} +``` + +**Purpose**: Returns the priority of this processor. Processors with higher priority are invoked before those with lower priority. The default is `0`. + +**Usage**: Allows ordering of processors when multiple might handle the same file. + +**Example**: + +```rust +fn priority(&self) -> i32 { + 10 +} +``` + +--- + +### Processor Implementation + +To create a custom processor, implement the `Processor` trait. Below is a complete example based on the test suite: + +```rust +use librawssg_error::Result; +use librawssg_fs::FileSystem; +use librawssg_handler::{Document, Metadata, Processor}; +use std::path::Path; + +struct MarkdownProcessor; + +impl Processor for MarkdownProcessor { + fn name(&self) -> &str { + "markdown" + } + + fn can_process(&self, relative_path: &Path, _original_path: &Path) -> bool { + relative_path.extension().and_then(|e| e.to_str()) == Some("md") + } + + fn process( + &self, + fs: &dyn FileSystem, + relative_path: &Path, + content_dir: &Path, + ) -> Result> { + // Read the source file + let source_path = content_dir.join(relative_path); + let content = fs.read_to_string(&source_path)?; + + // Parse front matter and body (simplified here) + let metadata = Metadata::new("Untitled", "")?; + let body = content; // In reality, you would render Markdown to HTML + + // Construct Document + let doc = Document::new( + metadata, + body, + relative_path.with_extension("html").to_string_lossy().to_string(), + relative_path.with_extension("index.html").to_string_lossy().into(), + source_path, + 1, + "page", + false, + )?; + + Ok(Some(doc)) + } +} +``` + +--- + +## Error Handling + +The library uses the `librawssg_error::Error` enum for all fallible operations. Relevant variants: + +- `Error::Validation(String)` – Used when a validation rule fails (e.g., empty URL, invalid depth). +- `Error::Processor(String)` – Used by processors to signal processing errors. +- Other variants may exist but are not directly used in this crate. + +The return type `Result` is an alias for `std::result::Result`. + +**Example of Validation Error**: + +```rust +let result = Document::new(meta, "body", "", "out", "src.md", 0, "page", false); +assert!(matches!( + result, + Err(librawssg_error::Error::Validation(ref msg)) if msg == "document url cannot be empty" +)); +``` + +--- + +## External Traits & Types + +### `FileSystem` Trait + +The `Processor::process` method takes a `&dyn FileSystem`. This trait is defined in `librawssg_fs` and provides a comprehensive set of file operations (read, write, create directories, walk, etc.). A typical implementation wraps `std::fs`, but for testing, mock implementations are often used. + +A minimal `FileSystem` implementation (used in tests) might implement all methods returning `io::Error::other("not implemented")` for those not needed. + +--- + +## Examples from Tests + +The test suite contains numerous examples that demonstrate correct usage and error conditions. Below are selected examples. + +### Creating a Valid Document + +```rust +use librawssg_handler::{Document, Metadata}; + +let metadata = Metadata::new("Title", "Description")?; +let doc = Document::new( + metadata, + "

Body

", + "blog/my-post.html", + "blog/my-post/index.html", + "content/blog/my-post.md", + 1, + "blog", + false, +)?; + +assert_eq!(doc.metadata.title, "Title"); +assert_eq!(doc.body, "

Body

"); +``` + +### Handling Invalid URL + +```rust +let result = Document::new( + Metadata::new("Title", "Desc")?, + "body", + "", // empty URL + "out", + "src.md", + 0, + "page", + false, +); +assert!(result.is_err()); +``` + +### Adding Taxonomy Terms + +```rust +let mut doc = /* ... */; +doc.add_taxonomy("categories", vec!["rust".to_string(), "ssg".to_string()]); +``` + +### Working with `extra` Metadata + +```rust +use serde_json::json; + +let mut meta = Metadata::new("Title", "Desc")?; +meta.insert_extra("key", "value"); +if let Some(v) = meta.get_extra("key") { + assert_eq!(v, &json!("value")); +} +``` + +### Processor Mock Example + +```rust +struct MockProcessor { /* ... */ } + +impl Processor for MockProcessor { + fn name(&self) -> &str { "mock" } + fn can_process(&self, _: &Path, _: &Path) -> bool { true } + fn process(&self, _fs: &dyn FileSystem, _rel: &Path, _cd: &Path) -> Result> { + Ok(Some(document)) + } +} +``` + +--- + +## Validation Rules Summary + +### `Metadata::new` + +- `title` must not be empty or contain only whitespace. + +### `Document::new` + +1. `url` must not be empty or contain only whitespace. +2. `output_path` must not be empty (as an OS string). +3. `source_path` must have a file name component. +4. `depth` must be ≤ 1000. + +Any violation results in an `Err(Error::Validation(...))`. + +--- + +## Testing Suite Overview + +The tests are organized into four files: + +1. **`document_tests.rs`** – Validates `Document` construction, field defaults, error cases, and methods (`relative_url`, `add_taxonomy`, `depth`, `with_list_items`). +2. **`metadata_tests.rs`** – Tests `Metadata` creation, validation, `is_draft`, `insert_extra`/`get_extra`, default values, and JSON serialization/deserialization round‑trip. +3. **`processor_tests.rs`** – Tests the `Processor` trait using mock implementations: `name`, `priority`, `can_process`, `process` returning `Some`, `None`, and error. +4. **`unit_tests.rs`** – Verifies that all public items are re‑exported at the crate root. + +All tests can serve as executable examples of the API usage. + +--- + +## Conclusion + +This documentation covers the public API of `librawssg_handler` in detail. The crate provides a flexible foundation for building static site generators by separating metadata handling (`Metadata`), document representation (`Document`), and pluggable processing logic (`Processor`). The validation rules ensure data integrity, and the use of traits like `FileSystem` enables testability. + +For further details, refer to the source code and the accompanying test suite. diff --git a/librawssg_templates/README.md b/librawssg_templates/README.md index 81b90c1..f21e43c 100644 --- a/librawssg_templates/README.md +++ b/librawssg_templates/README.md @@ -1,4 +1,656 @@ # librawssg_templates -Template rendering contracts for librawssg. -Defines the `Renderer` and `RenderContext` traits for pluggable template engines. +**Version**: 1.0.0 (implied) +**Crate name**: `librawssg_templates` +**Description**: Defines rendering abstractions for static site generators. Provides the `Renderer` and `RenderContext` traits, and an optional `TeraRenderer` implementation (when the `tera` feature is enabled) that integrates the Tera template engine. + +--- + +## Table of Contents + +1. [Overview](#overview) +2. [Modules and Features](#modules-and-features) +3. [Core Traits](#core-traits) + - [`RenderContext`](#trait-rendercontext) + - [Required Methods](#rendercontext-required-methods) + - [`Renderer`](#trait-renderer) + - [Required Method](#renderer-required-method) +4. [`TeraRenderer`](#struct-terarenderer) + - [Struct Definition](#struct-definition) + - [Constructor `new()`](#terarenderer-new) + - [Method `add_raw_template()`](#terarenderer-add_raw_template) + - [Method `add_template_file()`](#terarenderer-add_template_file) + - [Method `add_template_files_from_dir()`](#terarenderer-add_template_files_from_dir) + - [Method `load_templates_dir()`](#terarenderer-load_templates_dir) + - [Method `enable_autoescape()`](#terarenderer-enable_autoescape) + - [Method `render_str()`](#terarenderer-render_str) + - [Method `as_tera()`](#terarenderer-as_tera) + - [Method `as_tera_mut()`](#terarenderer-as_tera_mut) + - [Trait Implementations](#terarenderer-trait-implementations) + - [`Default`](#terarenderer-default) + - [`Renderer` for `TeraRenderer`](#terarenderer-renderer-impl) + - [`RenderContext` for `tera::Context`](#rendercontext-for-teracontext) +5. [Internal Helper Function](#internal-helper-function) +6. [Error Handling](#error-handling) +7. [Feature Gating](#feature-gating) +8. [Examples from Tests](#examples-from-tests) + - [Basic Rendering](#basic-rendering) + - [Loops, Filters, Conditions](#loops-filters-conditions) + - [File Loading](#file-loading) + - [Autoescaping](#autoescaping) + - [Template Inheritance and Macros](#template-inheritance-and-macros) + - [Context Downcasting](#context-downcasting) +9. [Testing Suite Overview](#testing-suite-overview) +10. [Conclusion](#conclusion) + +--- + +## Overview + +`librawssg_templates` provides a pluggable template rendering system. It abstracts the rendering process with two traits: + +- **`Renderer`** – Defines the `render` method that takes a template name and a context, returning a rendered string. +- **`RenderContext`** – An object‑safe trait that allows type erasure for context objects; specifically, it provides `as_any` and `as_mut_any` to downcast to concrete context types. + +The crate optionally includes a **`TeraRenderer`** implementation for the [Tera](https://tera.netlify.app/) template engine. This implementation is gated behind the `tera` feature flag. + +--- + +## Modules and Features + +The crate root (`lib.rs`) declares: + +```rust +pub mod renderer; +#[cfg(feature = "tera")] +pub mod tera_renderer; + +pub use renderer::{RenderContext, Renderer}; +#[cfg(feature = "tera")] +pub use tera_renderer::TeraRenderer; +``` + +- **`renderer`** – Always available; contains the two core traits. +- **`tera_renderer`** – Only compiled when the `tera` feature is enabled; contains `TeraRenderer`. +- Re‑exports at the crate root make the traits and `TeraRenderer` easy to import. + +The `tera` feature must be explicitly enabled in `Cargo.toml` to use `TeraRenderer`. Without it, the crate still provides the traits for custom renderer implementations. + +--- + +## Core Traits + +### Trait `RenderContext` + +```rust +pub trait RenderContext: Send + Sync { + fn as_any(&self) -> &dyn Any; + fn as_mut_any(&mut self) -> &mut dyn Any; +} +``` + +**Purpose**: Allows arbitrary context types to be passed to a `Renderer` as a trait object. The renderer can then downcast the `&dyn RenderContext` to the concrete context type it expects (e.g., `tera::Context`). This provides flexibility without requiring all renderers to accept a single concrete type. + +**Requirements**: + +- Implementors must be `Send + Sync` (thread‑safe). +- Must provide `as_any` and `as_mut_any` to expose the underlying `Any` reference. + +**Typical Implementation**: + +For any type `T`, you can implement: + +```rust +impl RenderContext for T { + fn as_any(&self) -> &dyn Any { + self + } + fn as_mut_any(&mut self) -> &mut dyn Any { + self + } +} +``` + +**Example** (from tests): + +```rust +struct MockContext; + +impl RenderContext for MockContext { + fn as_any(&self) -> &dyn Any { + self + } + fn as_mut_any(&mut self) -> &mut dyn Any { + self + } +} +``` + +--- + +### Trait `Renderer` + +```rust +pub trait Renderer: Send + Sync { + fn render(&self, template_name: &str, context: &dyn RenderContext) -> Result; +} +``` + +**Purpose**: Defines the rendering interface. A `Renderer` takes a template identifier (name) and a context object, and returns the rendered output as a `String`. + +**Parameters**: + +- `template_name`: A string identifying the template (e.g., `"index.html"`, `"blog/post.tera"`). +- `context`: A reference to an object implementing `RenderContext`. The renderer is expected to downcast this to the appropriate concrete context type. + +**Returns**: + +- `Ok(String)` containing the rendered output. +- `Err(librawssg_error::Error)` if rendering fails (e.g., template not found, invalid syntax, missing variable, or context type mismatch). + +**Note**: Implementors must be `Send + Sync`. + +**Example** (custom mock renderer): + +```rust +struct MockRenderer { output: String } + +impl Renderer for MockRenderer { + fn render(&self, _template_name: &str, _context: &dyn RenderContext) -> Result { + Ok(self.output.clone()) + } +} +``` + +--- + +## Struct `TeraRenderer` + +`TeraRenderer` is a wrapper around `tera::Tera`, providing convenient methods to load templates and render them using the `Renderer` trait. It is only available when the `tera` feature is enabled. + +### Struct Definition + +```rust +#[derive(Debug)] +pub struct TeraRenderer { + tera: tera::Tera, +} +``` + +The `tera` field is private; access is provided via `as_tera` and `as_tera_mut`. + +--- + +### `TeraRenderer::new` + +```rust +#[must_use] +pub fn new() -> Self +``` + +**Purpose**: Creates a new `TeraRenderer` with an empty Tera instance (`tera::Tera::default()`). + +**Returns**: A new `TeraRenderer`. + +**Example**: + +```rust +let renderer = TeraRenderer::new(); +``` + +--- + +### `TeraRenderer::add_raw_template` + +```rust +pub fn add_raw_template(&mut self, name: &str, content: &str) -> Result<()> +``` + +**Purpose**: Adds a template from a string, associating it with the given `name`. The template is parsed and stored internally. + +**Parameters**: + +- `name`: The template name (e.g., `"index.html"`, `"partial"`). +- `content`: The raw template source (e.g., `"Hello {{ name }}"`). + +**Returns**: + +- `Ok(())` if the template was added successfully. +- `Err(Error::Render)` if the template syntax is invalid (the underlying `tera::Error` is converted to a string and wrapped). + +**Behavior**: Calls `tera.add_raw_template(name, content)`. Template names must be unique; adding a duplicate name will replace the existing template. + +**Example**: + +```rust +renderer.add_raw_template("hello", "Hello {{ name }}")?; +``` + +--- + +### `TeraRenderer::add_template_file` + +```rust +pub fn add_template_file(&mut self, path: &Path) -> Result<()> +``` + +**Purpose**: Reads a template file from disk and adds it to the renderer. The template name is derived from the file name (including extension). + +**Parameters**: + +- `path`: Path to the template file. + +**Returns**: + +- `Ok(())` on success. +- `Err(Error::Io)` if the file cannot be read (wrapped as `Error::Io` with a message containing the original I/O error). +- `Err(Error::Render)` if the file name is not valid UTF‑8 or missing (unlikely). + +**Behavior**: + +1. Reads the file content using `std::fs::read_to_string`. On failure, maps to `Error::Io(std::io::Error::other(format!("{e}")))`. +2. Extracts the file name (the last component of the path) and converts it to a `&str`. If missing or non‑UTF‑8, returns `Error::Render("template file has no valid file name")`. +3. Calls `add_raw_template` with that file name as the template name. + +**Example**: + +```rust +renderer.add_template_file(Path::new("templates/index.html"))?; +// Template is registered as "index.html" +``` + +--- + +### `TeraRenderer::add_template_files_from_dir` + +```rust +pub fn add_template_files_from_dir(&mut self, dir: &Path) -> Result<()> +``` + +**Purpose**: Adds all files directly inside a directory as templates. This method is **not recursive**; it only considers files in the immediate directory. + +**Parameters**: + +- `dir`: Directory containing template files. + +**Returns**: + +- `Ok(())` if at least the directory is readable and processing completes. +- `Err(Error::Io)` on directory read failure. +- `Err(Error::Render)` if any individual file cannot be added. + +**Behavior**: + +1. Reads the directory entries using `std::fs::read_dir`. +2. For each entry: + - If the entry is a file, calls `add_template_file` with its path. + - If that returns an error, the method immediately returns the error (fail‑fast). +3. Non‑file entries (subdirectories, symlinks) are ignored. + +**Note**: Template names are the file names (including extensions). + +**Example**: + +```rust +renderer.add_template_files_from_dir(Path::new("templates/"))?; +// Adds all files in templates/ as templates with names like "base.tera", "index.html", etc. +``` + +--- + +### `TeraRenderer::load_templates_dir` + +```rust +pub fn load_templates_dir(&mut self, dir: &Path) -> Result<()> +``` + +**Purpose**: Recursively loads all template files from a directory tree. Template names are derived from the relative path (using forward slashes as separators), allowing nested template structures (e.g., `"sub/nested.tera"`). + +**Parameters**: + +- `dir`: Root directory to traverse. + +**Returns**: + +- `Ok(())` on success. +- `Err(Error::Io)` for filesystem errors during traversal or reading. +- `Err(Error::Render)` for invalid UTF‑8 paths or component issues. + +**Behavior**: + +1. Canonicalizes the input directory (using `dir.canonicalize()`) to ensure a stable base. +2. Walks the directory recursively using `walkdir::WalkDir`. Only files are processed. +3. For each file: + - Computes its path relative to the canonical directory using `strip_prefix`. + - Converts the relative path to a template name using the internal helper `rel_path_to_template_name` (which joins components with `/` and rejects non‑normal components). + - Reads the file content. + - Calls `add_raw_template` with the computed template name and content. + +**Note**: This method is similar to `add_template_files_from_dir` but recursive and with namespace‑like template names. + +**Example**: + +```rust +renderer.load_templates_dir(Path::new("templates"))?; +// If templates contains sub/child.tera, it can be referenced as "sub/child.tera" +``` + +--- + +### `TeraRenderer::enable_autoescape` + +```rust +pub fn enable_autoescape(&mut self) +``` + +**Purpose**: Turns on automatic escaping for HTML, HTM, and XML file extensions. This is a convenience method that calls `tera.autoescape_on(vec!["html", "htm", "xml"])`. + +**Parameters**: None. + +**Returns**: Nothing. + +**Behavior**: After calling this, templates with names ending in `.html`, `.htm`, or `.xml` will automatically escape variable output (HTML escaping). For other file extensions, autoescaping remains off. + +**Example**: + +```rust +renderer.enable_autoescape(); +renderer.add_raw_template("page.html", "{{ user_input }}")?; +// Rendering will escape HTML special characters in user_input +``` + +--- + +### `TeraRenderer::render_str` + +```rust +pub fn render_str(&self, template_str: &str, context: &dyn RenderContext) -> Result +``` + +**Purpose**: Renders a one‑off template string without registering it. This is useful for small, inline templates. + +**Parameters**: + +- `template_str`: The template source as a string. +- `context`: A `&dyn RenderContext` that must downcast to `tera::Context`. + +**Returns**: + +- `Ok(String)` with rendered output. +- `Err(Error::Render)` if the context is not a `tera::Context` or if rendering fails. + +**Behavior**: + +1. Attempts to downcast `context.as_any()` to `&tera::Context`. If the cast fails, returns `Error::Render("invalid context type for Tera")`. +2. Calls `tera::Tera::one_off(template_str, tera_ctx, true)`. The third argument `true` enables autoescaping for the one‑off render **by default**, regardless of the renderer's autoescape settings. + +**Note**: `render_str` always autoescapes (the `true` parameter forces autoescape). This may differ from `render` which respects the renderer's autoescape configuration. + +**Example**: + +```rust +let renderer = TeraRenderer::new(); +let mut ctx = tera::Context::new(); +ctx.insert("content", "bold"); +let output = renderer.render_str("{{ content }}", &ctx)?; +// Output is escaped: "<b>bold</b>" +``` + +--- + +### `TeraRenderer::as_tera` + +```rust +#[must_use] +pub const fn as_tera(&self) -> &tera::Tera +``` + +**Purpose**: Returns an immutable reference to the underlying `tera::Tera` instance. This allows advanced operations not directly exposed by `TeraRenderer`. + +**Returns**: `&tera::Tera`. + +**Example**: + +```rust +let tera = renderer.as_tera(); +// e.g., inspect registered templates +``` + +--- + +### `TeraRenderer::as_tera_mut` + +```rust +#[must_use] +pub const fn as_tera_mut(&mut self) -> &mut tera::Tera +``` + +**Purpose**: Returns a mutable reference to the underlying `tera::Tera` instance. Useful for direct manipulation, such as adding templates or changing settings. + +**Returns**: `&mut tera::Tera`. + +**Example**: + +```rust +let tera_mut = renderer.as_tera_mut(); +tera_mut.add_raw_template("direct", "Hello")?; +``` + +--- + +### Trait Implementations + +#### `TeraRenderer::Default` + +```rust +impl Default for TeraRenderer { + fn default() -> Self { + Self::new() + } +} +``` + +Allows creating a `TeraRenderer` with `TeraRenderer::default()`, equivalent to `new()`. + +#### `Renderer` for `TeraRenderer` + +```rust +impl Renderer for TeraRenderer { + fn render(&self, template_name: &str, context: &dyn RenderContext) -> Result { + let tera_ctx = context + .as_any() + .downcast_ref::() + .ok_or_else(|| Error::Render("invalid context type for Tera".into()))?; + self.tera + .render(template_name, tera_ctx) + .map_err(|e| Error::Render(e.to_string())) + } +} +``` + +- **`render` method**: + - Downcasts the context to `tera::Context`. + - Delegates to `tera.render(template_name, tera_ctx)`. + - Maps any error to `Error::Render`. + +#### `RenderContext` for `tera::Context` + +```rust +impl RenderContext for tera::Context { + fn as_any(&self) -> &dyn core::any::Any { + self + } + fn as_mut_any(&mut self) -> &mut dyn core::any::Any { + self + } +} +``` + +This implementation allows `tera::Context` to be used directly as a `RenderContext` when calling `render`. Users typically create a `tera::Context`, populate it, and pass `&ctx` to `render`. + +--- + +## Internal Helper Function + +`rel_path_to_template_name` (private) + +```rust +fn rel_path_to_template_name(rel_path: &Path) -> Result +``` + +**Purpose**: Converts a relative path (from `strip_prefix`) into a template name string using forward slashes as separators. It rejects paths with unusual components (prefixes, root, parent, or current directory). + +**Parameters**: + +- `rel_path`: A relative path (assumed to have been stripped of a base). + +**Returns**: + +- `Ok(String)` with the template name (e.g., `"sub/nested.tera"`). +- `Err(Error::Render)` if: + - A component is not `Normal` (e.g., contains `..` or `/` absolute parts). + - The path contains non‑UTF‑8 characters. + - The resulting name is empty. + +**Note**: This function is not public but is essential for `load_templates_dir`. + +--- + +## Error Handling + +All fallible methods in `TeraRenderer` return `librawssg_error::Result`. The errors originate from: + +- Filesystem operations → converted to `Error::Io` (with `std::io::Error::other` wrapper to preserve the original message). +- Tera template parsing/rendering → converted to `Error::Render` with the error message as string. +- Context type mismatch → `Error::Render("invalid context type for Tera")`. +- Invalid template file name or path component → `Error::Render`. + +The `Renderer` trait method also returns `Result`, allowing custom renderers to use the same error type. + +--- + +## Feature Gating + +- The core traits (`Renderer`, `RenderContext`) are always available. +- `TeraRenderer` and the `tera` integration are only compiled when the `tera` feature is enabled. +- Tests that use `TeraRenderer` are also gated with `#![cfg(feature = "tera")]`. + +To enable the feature, add to `Cargo.toml`: + +```toml +[dependencies] +librawssg_templates = { version = "...", features = ["tera"] } +``` + +--- + +## Examples from Tests + +The test suite (`tera_tests.rs` and `unit_tests.rs`) provides extensive examples. Below are selected snippets with explanations. + +### Basic Rendering + +```rust +let mut renderer = TeraRenderer::new(); +renderer.add_raw_template("simple", "{{ title }}")?; + +let mut ctx = tera::Context::new(); +ctx.insert("title", "Hello World"); +let output = renderer.render("simple", &ctx)?; +assert_eq!(output, "Hello World"); +``` + +### Loops, Filters, Conditions + +```rust +renderer.add_raw_template("loop", "{% for item in items %}{{ item }}{% if not loop.last %},{% endif %}{% endfor %}")?; +// context: items = ["a","b","c"] -> output "a,b,c" + +renderer.add_raw_template("filter", "{{ title | upper }}")?; +// context: title = "Hello" -> "HELLO" + +renderer.add_raw_template("condition", "{% if number > 40 %}high{% else %}low{% endif %}")?; +// context: number = 42 -> "high" +``` + +### File Loading + +```rust +// Add a single file +let file_path = dir.path().join("hello.tera"); +std::fs::write(&file_path, "{{ name }}")?; +renderer.add_template_file(&file_path)?; +// Template name is "hello.tera" + +// Add all files from a directory (non-recursive) +renderer.add_template_files_from_dir(dir.path())?; + +// Recursive loading with namespaced names +renderer.load_templates_dir(dir.path())?; +// If sub/nested.tera exists, use renderer.render("sub/nested.tera", &ctx) +``` + +### Autoescaping + +```rust +renderer.enable_autoescape(); +renderer.add_raw_template("esc.html", "{{ content }}")?; +let mut ctx = tera::Context::new(); +ctx.insert("content", ""); +let output = renderer.render("esc.html", &ctx)?; +// Output: <script>alert(1)</script> +``` + +Note: `render_str` always autoescapes: + +```rust +let output = renderer.render_str("{{ content }}", &ctx)?; +// Also escaped +``` + +### Template Inheritance and Macros + +```rust +renderer.add_raw_template("base", "{% block content %}Default{% endblock %}")?; +renderer.add_raw_template("child", "{% extends \"base\" %}{% block content %}Child content{% endblock %}")?; +// Rendering "child" yields "Child content" + +renderer.add_raw_template("macro", "{% macro hello(name) %}Hello, {{ name }}{% endmacro hello %}{{ self::hello(name=\"World\") }}")?; +// Rendering "macro" yields "Hello, World" +``` + +### Context Downcasting + +```rust +let ctx = sample_context(); +let dyn_ctx: &dyn RenderContext = &ctx; +assert!(dyn_ctx.as_any().is::()); +``` + +This shows how the `RenderContext` trait enables type erasure and safe downcasting. + +--- + +## Testing Suite Overview + +The crate contains two test files: + +- **`unit_tests.rs`** (always compiled): Tests the core traits using mock implementations, verifies re‑exports are available, and ensures the traits can be used without the `tera` feature. +- **`tera_tests.rs`** (compiled only with `tera` feature): Comprehensive tests for `TeraRenderer` including: + - Basic variable substitution, loops, filters, conditions. + - Error cases (missing template, missing variable, invalid syntax). + - Loading templates from files and directories (both non‑recursive and recursive). + - Autoescaping behavior. + - Template inheritance, macros, includes. + - Context downcasting. + - Access to underlying `Tera` via `as_tera` and `as_tera_mut`. + +The tests use `tempfile` for temporary directories and `walkdir` for directory traversal validation. + +--- + +## Conclusion + +`librawssg_templates` offers a flexible and extensible template rendering abstraction. The core `Renderer` and `RenderContext` traits allow any template engine to be integrated, while the built‑in `TeraRenderer` provides a powerful, full‑featured implementation for the Tera engine. With methods for loading templates from files or directories, autoescaping control, and direct access to the underlying engine, it covers the needs of most static site generators. + +The crate is designed with testability and thread‑safety in mind, and the comprehensive test suite serves as both documentation and validation. By enabling the `tera` feature, developers can immediately start rendering templates with minimal setup.