kodebeat / papers / ai

Deciding which models a machine can run before downloading one

A hardware-aware model, quantization and runtime advisor.

Reads the hardware it is running on, does the memory and bandwidth arithmetic, and recommends a model, a quantization and a runtime, with an estimated decode rate and the maximum context that will fit. It measures RAM bandwidth rather than assuming it, and can plan for hardware you do not own yet. For anyone who has downloaded a 40 GB model to find out whether it runs.

What is in it

  • The problem — what goes unanswered without it, and who notices first.
  • Why the obvious alternative falls short — stated plainly, including where it is the better choice.
  • How it works — the method, not a feature list.
  • Concrete use cases — with console output quoted from the repository, never reconstructed.
  • The methodology behind any number it emits — every term shown, so the figure survives a question.
  • What it deliberately does not do — the section most papers leave out.

Part of the AI infrastructure theme.

Get the PDF

One email with the download link, and this paper already selected. No follow-up sequence.

Send it to me