If anyone's looking for something a bit more minimal, I can't recommend hax [0] enough. No MCP, agent just gets a shell tool, simple config, whole thing's in C.
That looks interesting, built it (took seconds) and it is much smaller than Pi in terms of what it needs to work (when i installed Pi in a fresh Debian container it downloaded ~500MB of stuff, which isn't exactly what i had in mind when i read it is minimalistic :-P but it is a container so i didn't care much).
I'm just running it now in its own source with Qwen 3.8 27B and llama-server and asked it to analyze the code itself. I'm mainly curious to see how it handles context compaction during tasks (what Pi does is almost seamless and AFAICT it isn't anything particularly fancy so i'd expect Hax to do something similar) and i guess asking it to analyze a whole C codebase would help trigger that with a 131,072 context. Unfortunately it seems to be missing some "context usage" indicator while it does stuff (it shows context usage in the prompt but not while working), but i guess if it does manage to analyze the C code properly, i can ask it to add that :-P and see how it fares (from my use of Pi i'm positive Qwen 3.8 27B can do all that stuff, so it'd mainly be up to the harness).
EDIT: also i wonder if it works nicely if it is possible to convert Pi transcripts to Hax - i have a few "in progress" and i'd like to continue where i left from, though while both seem to use JSONL for the transcripts i'm not sure if they're compatible
EDIT2: hrm, it tried to use more than available tokens during a compaction and stopped there expecting me to increase the limit (i can, but what if i couldn't?) and restart the llama-server. Pi sometimes does hit it but it manages to recover by itself without requiring any input by me (or to increase llama-server's limit).
FWIW the compaction had failed again, so i used Pi with the same prompt, same model, same config to compare. The prompt was "check the current directory" followed by "analyze the code and give me a report in `CODE_REPORT.md` on how it works (also a brief report here). Make sure to update `CODE_REPORT.md` frequently (i.e. every time you analyze a file) to avoid context loss from context compaction".
Pi ended up hitting a compaction during a tool call and finished without issues, which is what i expected - in fact after giving the instruction i left to visit a relative since i expected it wouldn't need me to babysit it.
FWIW after it finished, i asked it the following:
---
Can you answer me the following questions about how Hax manages the context?
1. How does Hax handle running out of tokens? It does have some form of compaction, but if the compaction fails for some reason (e.g. the summary ends up needing more tokens) what does it do?
2. Can it handle cases where the context runs out of tokens during tool calls and if so, can it recover? How?
3. What happens with compaction if an LLM produces a few large responses or the LLM reads a few large files? Does the process ends up summarizing the entire (or all but one) conversation? Or is the entire conversation lost?
4. Are there any safeguards in place to avoid overwhelming the LLM? For example any file and tool output limits? If there is and any limits are reached, how does it handle them?
---
It went on and checked the code and, briefly the response (it was bigger but i don't want to repeat the entire thing):
1. It doesn't update the session on failure (last "good" session is kept), running out of context is treated like any other error without any automatic recovery and you're expected to fix it by hand (personally i'm not a fan of this).
2. Multiple tool calls are fine (there is a 85% threshold check to trigger compaction and a 50k tool result limit) but if a response and results jumps from below 85% to over 100% despite being under the 50k limit, it doesn't trigger any compaction and the next request is denied by the provider (i.e. llama-server). I have a feeling this is what i hit when i tried Hax with its own code.
3. There is no recovery from a user turn (prompt + response + tool result) that exceeds the window. FWIW this also seems to be the case with Pi.
4. It found a bunch of safeguards (tool output cap, caps in bytes and lines for the read tool, bash writes to a temp file and only a part of it is sent to the LLM, edit size cap, etc). AFAICT it is the same as Pi with one neat addition in that there is a default 2 minute timeout for bash calls (i've seen the LLM more than once run a command in Pi and end up stopping for 30+ minutes because the command wouldn't end).
For 3 i asked it a followup question: "About 3: AFAIK Pi (the harness you're on right now) does a "spit" summarization where the old messages are summarized up to a cutoff/split point and replaced with the summary while the newer messages after the cutoff/split point remain intact, which allow a mostly seamless transition between compactions. Does Hax do the same or something similar?"
The response was that, no, it doesn't, it summarizes the entire context. Which TBH is a bit of a dealbreaker for me since i often rely on this "seamless" continuity in my prompts and feels like the main reason why compactions feel like a non-issue with Pi.
Take the above with a grain of salt, i only checked the code for the compaction not using a split/cutoff point between older and recent messages, the rest are whatever Qwen 3.8 27B understood, but they do match my short empirical test. Also if my own understanding of the code is correct, it seems to be using the same system prompt for the summary as for regular/interactive use while AFAIK Pi uses a dedicated "you're an expert summarizer" (or something like that :-P) prompt. Not sure if it makes much or any difference, with LLMs being what they are, but TBH whatever Pi does works great IME.
On the other hand the idea of a self-contained native AI harness in C/C++ is enticing, especially one that doesn't have any network traffic outside of LLM-related stuff[0] and explicit user requests (Pi does try to autoupdate and has a separate opt-out telemetry beacon - both of which are disabled in different means, one via environment variable and another via a setting, which smells a bit like an dark pattern to me).
Anyway, this is the result of my findings about Hax. It is neat, but TBH the context handling is the main dealbreaker for me, especially since i'm often having the agent do something in the background (using a local LLM isn't exactly the speediest workflow) and do other stuff or leave the computer alone, so the last thing i want is to babysit the agent for errors. Pi's split summarization and context overflow handling seem to work much better.
For now i'll probably stick with Pi (i have autoupdates and telemetry disabled and i hope there isn't any other hidden snitch in place) and perhaps at some point i'll do the NIH thing and make yet another agent myself :-P
[0] well, it does attempt to autoconnect to a potentially running llama-server in localhost without being explicitly told to do so (Pi wants explicit configuration) but meh
I use it with a local Qwen3.8 27B on llama-server, but I don't know what it means for it exceed any window.
you can use /slots on the llama-server if you want to get more up-to-date details on session token use.
I always run models in their default context, which is 262144 for Qwen3.8 27B. I've run sessions in hax where it hit that cap multiple times and compressed the context down to 15% and continued with no problem.
> also i wonder if it works nicely if it is possible to convert Pi transcripts to Hax - i have a few "in progress" and i'd like to continue where i left from, though while both seem to use JSONL for the transcripts i'm not sure if they're compatible
I’ll sometimes ask the new agent to look at the previous transcript file.
ltrg · · focus · HN ↗
[0] <a href="https://usehax.dev/" rel="nofollow">https://usehax.dev/
badsectoracula · · focus · HN ↗
I'm just running it now in its own source with Qwen 3.8 27B and llama-server and asked it to analyze the code itself. I'm mainly curious to see how it handles context compaction during tasks (what Pi does is almost seamless and AFAICT it isn't anything particularly fancy so i'd expect Hax to do something similar) and i guess asking it to analyze a whole C codebase would help trigger that with a 131,072 context. Unfortunately it seems to be missing some "context usage" indicator while it does stuff (it shows context usage in the prompt but not while working), but i guess if it does manage to analyze the C code properly, i can ask it to add that :-P and see how it fares (from my use of Pi i'm positive Qwen 3.8 27B can do all that stuff, so it'd mainly be up to the harness).
EDIT: also i wonder if it works nicely if it is possible to convert Pi transcripts to Hax - i have a few "in progress" and i'd like to continue where i left from, though while both seem to use JSONL for the transcripts i'm not sure if they're compatible
EDIT2: hrm, it tried to use more than available tokens during a compaction and stopped there expecting me to increase the limit (i can, but what if i couldn't?) and restart the llama-server. Pi sometimes does hit it but it manages to recover by itself without requiring any input by me (or to increase llama-server's limit).
badsectoracula · · focus · HN ↗
Pi ended up hitting a compaction during a tool call and finished without issues, which is what i expected - in fact after giving the instruction i left to visit a relative since i expected it wouldn't need me to babysit it.
FWIW after it finished, i asked it the following:
---
Can you answer me the following questions about how Hax manages the context?
1. How does Hax handle running out of tokens? It does have some form of compaction, but if the compaction fails for some reason (e.g. the summary ends up needing more tokens) what does it do?
2. Can it handle cases where the context runs out of tokens during tool calls and if so, can it recover? How?
3. What happens with compaction if an LLM produces a few large responses or the LLM reads a few large files? Does the process ends up summarizing the entire (or all but one) conversation? Or is the entire conversation lost?
4. Are there any safeguards in place to avoid overwhelming the LLM? For example any file and tool output limits? If there is and any limits are reached, how does it handle them?
---
It went on and checked the code and, briefly the response (it was bigger but i don't want to repeat the entire thing):
1. It doesn't update the session on failure (last "good" session is kept), running out of context is treated like any other error without any automatic recovery and you're expected to fix it by hand (personally i'm not a fan of this).
2. Multiple tool calls are fine (there is a 85% threshold check to trigger compaction and a 50k tool result limit) but if a response and results jumps from below 85% to over 100% despite being under the 50k limit, it doesn't trigger any compaction and the next request is denied by the provider (i.e. llama-server). I have a feeling this is what i hit when i tried Hax with its own code.
3. There is no recovery from a user turn (prompt + response + tool result) that exceeds the window. FWIW this also seems to be the case with Pi.
4. It found a bunch of safeguards (tool output cap, caps in bytes and lines for the read tool, bash writes to a temp file and only a part of it is sent to the LLM, edit size cap, etc). AFAICT it is the same as Pi with one neat addition in that there is a default 2 minute timeout for bash calls (i've seen the LLM more than once run a command in Pi and end up stopping for 30+ minutes because the command wouldn't end).
For 3 i asked it a followup question: "About 3: AFAIK Pi (the harness you're on right now) does a "spit" summarization where the old messages are summarized up to a cutoff/split point and replaced with the summary while the newer messages after the cutoff/split point remain intact, which allow a mostly seamless transition between compactions. Does Hax do the same or something similar?"
The response was that, no, it doesn't, it summarizes the entire context. Which TBH is a bit of a dealbreaker for me since i often rely on this "seamless" continuity in my prompts and feels like the main reason why compactions feel like a non-issue with Pi.
Take the above with a grain of salt, i only checked the code for the compaction not using a split/cutoff point between older and recent messages, the rest are whatever Qwen 3.8 27B understood, but they do match my short empirical test. Also if my own understanding of the code is correct, it seems to be using the same system prompt for the summary as for regular/interactive use while AFAIK Pi uses a dedicated "you're an expert summarizer" (or something like that :-P) prompt. Not sure if it makes much or any difference, with LLMs being what they are, but TBH whatever Pi does works great IME.
On the other hand the idea of a self-contained native AI harness in C/C++ is enticing, especially one that doesn't have any network traffic outside of LLM-related stuff[0] and explicit user requests (Pi does try to autoupdate and has a separate opt-out telemetry beacon - both of which are disabled in different means, one via environment variable and another via a setting, which smells a bit like an dark pattern to me).
Anyway, this is the result of my findings about Hax. It is neat, but TBH the context handling is the main dealbreaker for me, especially since i'm often having the agent do something in the background (using a local LLM isn't exactly the speediest workflow) and do other stuff or leave the computer alone, so the last thing i want is to babysit the agent for errors. Pi's split summarization and context overflow handling seem to work much better.
For now i'll probably stick with Pi (i have autoupdates and telemetry disabled and i hope there isn't any other hidden snitch in place) and perhaps at some point i'll do the NIH thing and make yet another agent myself :-P
[0] well, it does attempt to autoconnect to a potentially running llama-server in localhost without being explicitly told to do so (Pi wants explicit configuration) but meh
tingletech · · focus · HN ↗
you can use /slots on the llama-server if you want to get more up-to-date details on session token use.
I always run models in their default context, which is 262144 for Qwen3.8 27B. I've run sessions in hax where it hit that cap multiple times and compressed the context down to 15% and continued with no problem.
collinmanderson · · focus · HN ↗
I’ll sometimes ask the new agent to look at the previous transcript file.