For some reason reasoning effort low and medium used similar numbers of tokens, and xhigh used less than high. I think I need to try without OpenRouter in the middle.
UPDATE: I tried again with the xAI API directly: <a href="https://tools.simonwillison.net/markdown-svg-renderer?url=https%3A%2F%2Fgist.github.com%2Fsimonw%2F5a1819a2bd24bb642f38c4bc6733090f" rel="nofollow">https://tools.simonwillison.net/markdown-svg-renderer?url=ht... - not a great deal of difference between reasoning levels, and this time xhigh and low used the same number of reasoning tokens for some reason.
For comparison here's a fresh run against Grok 4.6: <a href="https://tools.simonwillison.net/markdown-svg-renderer?url=https%3A%2F%2Fgist.github.com%2Fsimonw%2Ffedc404b9aa8e6fca31d59c898dabba0" rel="nofollow">https://tools.simonwillison.net/markdown-svg-renderer?url=ht...
Are there good tools for doing context audits? I feel I have no good way to visualize what a new session is getting by default in a given repo without crawling through every potentially included markdown file
I tried in Cursor and see a lot of improvements over Grok 4.6 svgs. The AA numbers indicate it's not very token efficient though: <a href="https://artificialanalysis.ai/agents/coding-agents?agents=codex-deepseek-v4-pro-0813-max%2Ccodex-gpt-6-astra-max-reasoning-effort-max%2Cclaude-code-opus-5-max%2Ccodex-gpt-5-6-sol-max-reasoning-effort-max%2Cdevin-fusion-cli-claude-fable-5-1-xhigh-swe-2-medium%2Cclaude-code-qwen3-8-max%2Cdevin-fusion-cli-gpt-6-astra-xhigh-swe-2-medium%2Cantigravity-sdk-gemini-3-8-flash-high%2Cmuse-code-muse-spark-1-3-max%2Ckimi-code-cli-kimi-k3%2Cclaude-code-fable-5-1-max-with-fallback%2Copencode-glm-5-3-reasoning-effort-max%2Cgrok-build-grok-4-7-xhigh%2Cgrok-build-grok-4-6-xhigh#coding-agents-token-usage-chart-tabs" rel="nofollow">https://artificialanalysis.ai/agents/coding-agents?agents=co...
exactly, the benchmark just needs to be downvoted into oblivion each times it's posted. The outcome is not deterministic and the model needs to determine what level of detail is appropriate for an svg. There is no wrong answer to this unless it's obviously un-Pelican-like.
Yeah, he had opinions: <a href="https://twitter.com/elonmusk/status/2023833496804839808" rel="nofollow">https://twitter.com/elonmusk/status/2023833496804839808
simonw · · focus · HN ↗
Here's reasoning level high: <a href="https://tools.simonwillison.net/markdown-svg-renderer?url=https%3A%2F%2Fgist.github.com%2Fsimonw%2F8ab126bda2b384264b3ad931e3ebb8b4" rel="nofollow">https://tools.simonwillison.net/markdown-svg-renderer?url=ht...
For some reason reasoning effort low and medium used similar numbers of tokens, and xhigh used less than high. I think I need to try without OpenRouter in the middle.
UPDATE: I tried again with the xAI API directly: <a href="https://tools.simonwillison.net/markdown-svg-renderer?url=https%3A%2F%2Fgist.github.com%2Fsimonw%2F5a1819a2bd24bb642f38c4bc6733090f" rel="nofollow">https://tools.simonwillison.net/markdown-svg-renderer?url=ht... - not a great deal of difference between reasoning levels, and this time xhigh and low used the same number of reasoning tokens for some reason.
For comparison here's a fresh run against Grok 4.6: <a href="https://tools.simonwillison.net/markdown-svg-renderer?url=https%3A%2F%2Fgist.github.com%2Fsimonw%2Ffedc404b9aa8e6fca31d59c898dabba0" rel="nofollow">https://tools.simonwillison.net/markdown-svg-renderer?url=ht...
MattDamonSpace · · focus · HN ↗
datsci_est_2015 · · focus · HN ↗
kiliancs · · focus · HN ↗
forgot-my-pw · · focus · HN ↗
TomGarden · · focus · HN ↗
Mashimo · · focus · HN ↗
athrowaway3z · · focus · HN ↗
It used to be a mess in various interesting ways. Now, almost every big release can draw something perfectly functional.
So the question - without a correct answer - given the prompt "Generate an SVG of a pelican riding a bicycle":
Does the user want the least lines of code to make it functional, or the best looking version?
wolttam · · focus · HN ↗
peder · · focus · HN ↗
paimapi · · focus · HN ↗
daveguy · · focus · HN ↗
simonw · · focus · HN ↗
nicolamanzini · · focus · HN ↗
[dead]
danappelxx · · focus · HN ↗
iankp · · focus · HN ↗
mrtesthah · · focus · HN ↗
[dead]
jfoster · · focus · HN ↗
Has a shadow
Better shaped beak
Leg position more realistic for bicycle riding
Better feathers