llamaperf

firish/webfetch

Runs locallyNo key needed

Own your LLM's web search: a local search->fetch->rank pipeline that replaces hosted web-search tools. Measured: matches hosted accuracy at 66% lower cost and up to 88% fewer tokens, plus a precision-tuned semantic caching with query-dependant TTL that no API offers.

View on GitHub

Stars
58
Gained, 7 days
0
Gained, 30 days
n/a
Reddit posts
2

Python · MIT · last push 2026-07-23 · created 2026-04-20

GitHub stars, last 90 days

2026-09-26: 562026-10-07: 58

Does it work with local models?

Reports from people who ran it with a model on their own hardware. Tool calling depends on the model as much as on the server, so say which model and quant you used.

No reports yet.

Add your report

Sign in to say whether it worked for you.

Reddit posts linking it

In the official MCP registry

  • io.github.firish/webfetchv0.1.3Runs locallyNo key needed

    Self-hosted web search for LLM agents: search -> fetch -> rank pipeline with semantic caching

    Packages: pypi webfetch-llm (stdio)