Skip to content
#

dynamic-context-pruning

Here is 1 public repository matching this topic...

Drop-in FastAPI proxy for llama.cpp, Ollama, vLLM and similar backends. Automatically prunes, summarizes, and extracts the most relevant context from large inputs using advanced strategies (readagent, rlm, embeddings) so your local models answer accurately without cloud services.

  • Updated Sep 4, 2026
  • Python

Add this topic to your repo

To associate your repository with the dynamic-context-pruning topic, visit your repo's landing page and select "manage topics."

Learn more