The llama.cpp guys don't get nearly enough credit for their work. Though the quality of the codebase is dropping over time, it is still quite high compared to most of the alternatives, and it is still one of the most stable ways to run a large variety of models.
Definitely worth looking at if you have only a single 5090 is ninfer, and various hardware specific forks (3090, 4090).
HoldOnAMinute · · focus · HN ↗
pydry · · focus · HN ↗
jminnl · · focus · HN ↗
Definitely worth looking at if you have only a single 5090 is ninfer, and various hardware specific forks (3090, 4090).