I encapsulate Hugging Face model loading and inference within my Small_LLM_Model class. On the first run, initializing a model like "Qwen/Qwen3-0.6B" automatically fetches the complete model asset bundle from the Hugging Face Hub and caches it locally (~/.cache/huggingface/).
llm = Small_LLM_Model(model_name="Qwen/Qwen3-0.6B")
prompt = "Is 4 an even number?"
input_ids = llm._tokenizer.encode(prompt, add_special_tokens=False)
logits = llm.get_logits_from_input_ids(input_ids)Loading a causal language model requires three distinct categories of files working in tandem:
vocab.json: Map of sub-word token strings to integer IDs (e.g.,"hello" -> 31373).merges.txt: Ranked Byte-Pair Encoding (BPE) rules for combining single characters into sub-words.tokenizer.json: A unified, portable configuration combining vocabulary maps and merge rules into a single structure.tokenizer_config.json: Metadata defining tokenizer behavior, special control tokens (<eos>,<pad>), and padding sides.
Describes the exact neural network architecture and hyperparameters, allowing transformers to reconstruct the uninitialized PyTorch model layers in memory:
- Hidden layer depth (
n_layer) - Attention heads per layer (
n_head) - Embedding dimensions (
hidden_size) - Maximum context window (
n_positions)
Contains the trained parameter matrices (floating-point weights and biases) for embeddings, attention mechanisms, and feed-forward networks:
- Format: Uses
.safetensorsinstead of legacy pickle.binfiles to prevent arbitrary code execution vulnerabilities and accelerate zero-copy memory mapping. - Memory Footprint: Represents the bulk of the download size (typically several gigabytes).
When my pipeline processes a prompt, execution follows a 5-step sequence:
[ Raw Text ] ──> Tokenizer ──> [ 2D Tensor ] ──> Model (safetensors) ──> [ Logits ] ──> Output Text
- Asset Fetching: Reads cached weights and configurations from the local filesystem.
- Tokenizer Initialization:
AutoTokenizerloadstokenizer_config.json,vocab.json, andmerges.txtinto memory. - Architecture Construction:
AutoModelForCausalLMreadsconfig.jsonand builds the uninitialized PyTorch network structure. - Weight Injection: Parameters from
model.safetensorsare assigned directly to network layer weights. - Inference & Tensor Processing: Encodings produce a 2D batch tensor (
(batch_size, sequence_length)). To extract a flat 1D list of integer IDs for custom processing, I unpack the first row via.tolist()[0]:
# Converts 2D PyTorch tensor [[id1, id2, ...]] to flat Python list [id1, id2, ...]
input_ids_list = llm._encode(prompt).tolist()[0]