Browsing: Runs

You load a 7B parameter model onto a GPU. It fits comfortably. The first few prompts return without issue. Ten messages into the conversation, though, an out-of-memory error kills the process. This article breaks down exactly where every byte of VRAM goes during local LLM inference, provides a working Python calculator for estimating peak memory…