Skip to content
View SuperMarioYL's full-sized avatar
🎯
Focusing
🎯
Focusing

Block or report SuperMarioYL

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
SuperMarioYL/README.md
EN  ⇄  中文
Leo — AI systems, made to run in production.

I work across AI agents, cloud-native AI and inference acceleration: Agent Loop / RSI, disaggregated P/D deployments, and tiered KV cache reuse.

System architecture

The agent layer uses model services; the inference layer connects engines to tiered storage for KV cache reuse; the cloud-native layer provides routing, P/D deployment and compute resources.

The agent layer uses model services; the inference layer connects engines to tiered storage for KV cache reuse; the cloud-native layer provides routing, P/D deployment and compute resources.

Capabilities

Agent Loop and RSI, disaggregated Prefill/Decode deployment, and vLLM Connector offload with tiered KV storage.

Tech stack

Tech stack

Parallel tracks

Parallel tracks

Selected work


Happy to talk tech and build together.

blog email github views

Pinned Loading

  1. trouve trouve Public

    trouve : A built-in integrated service discovery, service registration, and service forwarding general component for Spring projects

    Java 31 9

  2. Bison Bison Public

    Enterprise GPU Resource Billing & Multi-Tenant Management Platform 企业级 GPU 资源计费与多租户管理平台

    TypeScript 7

  3. inference-cookbook inference-cookbook Public

    inference cookbook / inference 框架原理解析

    HTML 10 2

  4. tokensched tokensched Public

    Simulate value-aware token allocation across a task tree, choosing model tiers or preempting tasks within a declared budget.

    Go 3

  5. ModelEngine-Group/unified-cache-management ModelEngine-Group/unified-cache-management Public

    Persist and reuse KV Cache to speedup your LLM.

    C++ 328 112

  6. vllm-project/vllm vllm-project/vllm Public

    A high-throughput and memory-efficient inference and serving engine for LLMs

    Python 91.4k 22k