
Shoehorn
situationalQuantize any language model to fit your machine's actual VRAM, then serve it locally.
Visit notactuallytreyanastasio.github.io →> **Shoehorn** — situational > Verified 2026-08-19 > Limitation: Requires a working llama.cpp setup and some CLI comfort; the web UI still shells the same pipeline. > > https://noemium.com/tools/shoehorn/Add to kitOpen in NoemiumFact-check this catalog entry for Shoehorn (dev-infra). Verdict: situational Verdict text: A practical answer to "will this model run on my GPU?" Shoehorn solves a per-tensor mixed-precision assignment from your real memory budget instead of forcing a preset quantization, squeezing out quality other presets waste. Free, local, and genuinely useful for anyone running LLMs on consumer hardware. Limitations: - Requires a working llama.cpp setup and some CLI comfort; the web UI still shells the same pipeline. - Windows support is compile-tested but not yet field-tested. - Auto-generated imatrices are skipped on Windows, so you need repos that publish one. Receipts: https://github.com/notactuallytreyanastasio/shoehorn Last verified: 2026-08-19 Find outdated prices, wrong claims, or missing limitations. Cite primary sources (docs, pricing pages, benchmarks, model cards). Do not use marketing pages.[](https://noemium.com/tools/shoehorn/)
ledger entry · observed by @whysanesanders
A practical answer to "will this model run on my GPU?" Shoehorn solves a per-tensor mixed-precision assignment from your real memory budget instead of forcing a preset quantization, squeezing out quality other presets waste. Free, local, and genuinely useful for anyone running LLMs on consumer hardware.
Strengths
- Quantizes any language model to fit your machine's actual VRAM budget.
- Per-tensor mixed-precision assignment squeezes out quality other presets waste.
- Free and local for running LLMs on consumer hardware.
Known limitations
- Requires a working llama.cpp setup and some CLI comfort; the web UI still shells the same pipeline.
- Windows support is compile-tested but not yet field-tested.
- Auto-generated imatrices are skipped on Windows, so you need repos that publish one.
Use it when
- Local LLM inference on consumer GPUs with tight VRAM constraints.
- Scenarios where preset quantization throws away usable quality.
Skip it when
- You are not comfortable with llama.cpp and command-line workflows.
- You run Windows and need auto-generated imatrices for every model.
- A one-click web GUI is a hard requirement.
Buy-check
- ·Price — free, open source; bring your own compute and API keys as of 2026-08-19. Open the pricing page, not a screenshot.
- ·Free tier — Free tier exists — burn it first.
- ·Cancel fallout — What happens to your data/projects on cancel — check the ToS, not the FAQ.
- ·Seat math — Per-seat pricing multiplies quietly. Count seats before the annual plan.
- ·Who skips it — Also skipped when: You are not comfortable with llama.cpp and command-line workflows.; You run Windows and need auto-generated imatrices for every model..
Price trailchangelog →
No price move in this repo yet.
Facts
- pricing
- free1
- price note
- open source; bring your own compute and API keys
- free tier
- yes
- open source
- yes
- api
- no
- self-host
- yes
- category
- dev-infra
- evidence
- source-verified
Data: trains on inputs unknown · local processing yes
Sources
No source, no number.
Similar joball alternatives →
| tool | verdict | price | oss | self-host |
|---|---|---|---|---|
| Shoehorn | situational | free | yes | yes |
| Modal | ship | freemium | no | no |
| OpenRouter | ship | freemium | no | no |
| Replicate | ship | paid | no | no |
Facts from the catalog files — not scores you can buy. Quality and speed stay in the briefing, not in a fake 1–5 grid.
Modal
Serverless GPU and compute platform for running ML and AI workloads.
APIFREE TIER
freemiumfree tier ($30/mo compute); pay-as-you-go beyond
OpenRouteranchor
One API for every major model — unified billing, fallbacks and routing.
APIFREE TIER
freemiumpay-per-token pass-through plus a small fee
Replicate
Run any community model via API — pay per second, no GPU wrangling.
API
paidpay per run; image models from ~$0.003/image