Try “car wash”, “subscription box”, “Austin” · Esc to close

Shoehorn

Quantize AI models to run locally on your machine with a simple GUI

Developer tool / API SaaS & software Show HN · launch post · ▲ 99

Visit site

notactuallytreyanastasio.github.io

What it does

Shoehorn is a tool that quantizes AI language models to run locally on a user's machine by optimizing memory usage. It takes a BF16 GGUF model and uses an imatrix to assign mixed-precision quantization on a per-tensor basis, fitting the model precisely into available VRAM with minimal waste. The tool includes a GUI for discovering models and setting up inference, backed by llama.cpp as the inference engine.

Who it is for

Shoehorn targets developers and users who want to run language models locally on consumer hardware—from MacBooks to Linux systems with GPUs to Windows machines. The tool appeals to anyone constrained by specific memory budgets who wants to maximize model quality within those constraints rather than choosing from preset quantization options.

Pricing

The site does not show prices.

How it stands out

Most quantization approaches offer preset options (like 4-bit, 8-bit quantization) that either waste memory or fail to fit at runtime. Shoehorn inverts the problem: it starts with the user's actual hardware memory, subtracts inference overhead, then solves for optimal per-tensor precision assignments that use nearly 100% of the available budget. The homepage example shows 99.998% utilization with only 13 KB slack. The tool includes a browser-based hardware selector that scans Hugging Face's most-downloaded models to show which ones fit a given memory configuration, ranked by achievable quality. Installation is straightforward for macOS users via Homebrew, with Linux and Windows support also available. The interface is minimal: one button to quantize and prepare a model, then immediate access to a local chat interface.

What a founder should check

A founder building a competing product should verify: (1) whether existing quantization tools and preset libraries adequately serve users or if precise per-tensor optimization offers a meaningful quality advantage that justifies complexity; (2) the switching costs for users already using llama.cpp with standard quantizations—what gains in usable model quality would justify retraining or re-quantizing existing deployments; and (3) whether the moat is sustainable given that quantization logic itself is becoming commoditized and that llama.cpp and similar inference engines are open-source projects that could adopt similar optimization strategies.

Thinking of building something like this?

Every launch here is a competitor to somebody's idea. If yours is close, check it against the market before you build: the Full Check names the rivals, the prices and the gaps.

Check an idea like this

More developer tool / api launches

All

Wispbit

Linter that enforces codebase standards with AI coding agents.

Developer tool / API SaaS & softwareShow HN ▲ 31

OnlyJPG

Private browser-based converter for any image format to JPG.

Developer tool / API SaaS & softwareShow HN ▲ 64

Duck-UI

Browser-based SQL IDE for DuckDB running entirely in WebAssembly.

Developer tool / API SaaS & softwareShow HN ▲ 213

Checked ideas in SaaS & software