Skip to content

Catalogdev-infra

Shoehorn

situationalBlueshift — gaining momentum

Quantize any language model to fit your machine's actual VRAM, then serve it locally.

Visit notactuallytreyanastasio.github.io →Add to kitOpen in Noemium

ledger entry · observed by @whysanesanders

A practical answer to "will this model run on my GPU?" Shoehorn solves a per-tensor mixed-precision assignment from your real memory budget instead of forcing a preset quantization, squeezing out quality other presets waste. Free, local, and genuinely useful for anyone running LLMs on consumer hardware.

Strengths

  • Quantizes any language model to fit your machine's actual VRAM budget.
  • Per-tensor mixed-precision assignment squeezes out quality other presets waste.
  • Free and local for running LLMs on consumer hardware.

Known limitations

  • Requires a working llama.cpp setup and some CLI comfort; the web UI still shells the same pipeline.
  • Windows support is compile-tested but not yet field-tested.
  • Auto-generated imatrices are skipped on Windows, so you need repos that publish one.

Use it when

  • Local LLM inference on consumer GPUs with tight VRAM constraints.
  • Scenarios where preset quantization throws away usable quality.

Skip it when

  • You are not comfortable with llama.cpp and command-line workflows.
  • You run Windows and need auto-generated imatrices for every model.
  • A one-click web GUI is a hard requirement.

Buy-check

  • ·Price — free, open source; bring your own compute and API keys as of 2026-08-19. Open the pricing page, not a screenshot.
  • ·Free tier — Free tier exists — burn it first.
  • ·Cancel fallout — What happens to your data/projects on cancel — check the ToS, not the FAQ.
  • ·Seat math — Per-seat pricing multiplies quietly. Count seats before the annual plan.
  • ·Who skips it — Also skipped when: You are not comfortable with llama.cpp and command-line workflows.; You run Windows and need auto-generated imatrices for every model..

Price trailchangelog →

No price move in this repo yet.

Facts

pricing
free1
price note
open source; bring your own compute and API keys
free tier
yes
open source
yes
api
no
self-host
yes
category
dev-infra
evidence
source-verified

Data: trains on inputs unknown · local processing yes

Sources

  1. 1github.com/notactuallytreyanastasio/shoehorn[snapshot]

No source, no number.

Edit this page on GitHub

Similar joball alternatives →

toolverdictpriceossself-host
Shoehornsituationalfreeyesyes
Modalshipfreemiumnono
OpenRoutershipfreemiumnono
Replicateshippaidnono

Facts from the catalog files — not scores you can buy. Quality and speed stay in the briefing, not in a fake 1–5 grid.

Modal
dev-infraship

Modal

Serverless GPU and compute platform for running ML and AI workloads.

APIFREE TIER

freemiumfree tier ($30/mo compute); pay-as-you-go beyond

Blueshift — gaining momentum
OpenRouter
dev-infraship

OpenRouteranchor

One API for every major model — unified billing, fallbacks and routing.

APIFREE TIER

freemiumpay-per-token pass-through plus a small fee

Blueshift — gaining momentum
Replicate
dev-infraship

Replicate

Run any community model via API — pay per second, no GPU wrangling.

API

paidpay per run; image models from ~$0.003/image

Steady momentum