Brood War Bench
Thread
Loading the complete thread in the background. This saved snapshot is available now. Refresh
Unofficial Hacker News client; not affiliated with Y Combinator.
Brood War Bench
Loading the complete thread in the background. This saved snapshot is available now. Refresh
Unofficial Hacker News client; not affiliated with Y Combinator.
benswerd · · focus · HN ↗
conorcleary · · focus · HN ↗
indrora · · focus · HN ↗
bee_rider · · focus · HN ↗
stackghost · · focus · HN ↗
frutiger · · focus · HN ↗
benswerd · · focus · HN ↗
duskwuff · · focus · HN ↗
tekla · · focus · HN ↗
xmcp123 · · focus · HN ↗
duskwuff · · focus · HN ↗
xmcp123 · · focus · HN ↗
duskwuff · · focus · HN ↗
rrr_oh_man · · focus · HN ↗
faeyanpiraat · · focus · HN ↗
dmichulke · · focus · HN ↗
If you want 6 carriers at the same time, build 6 starports.
If you want 6 ultralisks or 6 mutas at the same time, build 2 hatches.
iammjm · · focus · HN ↗
malfist · · focus · HN ↗
shard972 · · focus · HN ↗
[dead]
tyre · · focus · HN ↗
Not sure they’re the best option for raiding, but as a high-level orchestrator for choosing content, that sounds pretty great.
stymaar · · focus · HN ↗
Avicebron · · focus · HN ↗
Game_Ender · · focus · HN ↗
benswerd · · focus · HN ↗
For agent harness I did Claude Code, Codex, Grok Build. This was primarily a cost driven decision — I have a lot of free tokens and I didn't want to pay API prices for this.
For game harness I used minimal BW-API issue command and get observation apis as tools. I felt this was the most fair way to do it on my small scale.
In the future I would like to integrate code mode and multiple games/I think if it was a best of 5 where each agent could learn from its past games and build its own automations over time that would be much more interesting.
usef- · · focus · HN ↗
Or possibly even whether they can learn from a game themselves. "Analyse your game for your failures" -> Then give a fresh agent of the same model that "learnings" doc for the next match. Do the rankings change over time, if models can write instructions for future selves that actually help?
benswerd · · focus · HN ↗
pelagicAustral · · focus · HN ↗
I love StarCraft. I started playing it right from the beginning, most of my friends right now are from that era. I literally met people that have spread to almost every continent when I was in my early teens. We played at internet cafes and did not have access to the internet, that was priced differently...
I miss those days so much.
Everybody was from a different background back then, and nobody was anything other than a guy that plays StaCraft at the cybercafe... And now, we are in our 40's and I know Math teachers, history teachers, oil rig operators, software programmers, professional gamers, lawyers and more... hahah So crazy to think about it... and I know them, we talk, what a world.
sidewndr46 · · focus · HN ↗
benswerd · · focus · HN ↗
My first time playing StarCraft was at summer camp around a decade after it came out.
All the smartest people played it so I wanted to too. Great decision, I have been continually impressed with the people who StarCraft introduced me to.
nemo1618 · · focus · HN ↗
benswerd · · focus · HN ↗
PorciiVorbesc · · focus · HN ↗
jpgvm · · focus · HN ↗
oddsockmachine · · focus · HN ↗
pelagicAustral · · focus · HN ↗
gary17the · · focus · HN ↗
sizzle · · focus · HN ↗
snicky · · focus · HN ↗
ddlsmurf · · focus · HN ↗
layoric · · focus · HN ↗
stavros · · focus · HN ↗
pelagicAustral · · focus · HN ↗
Forgeties79 · · focus · HN ↗
sizzle · · focus · HN ↗
egeozcan · · focus · HN ↗
[deleted] · · focus · HN ↗
[deleted]
kentosi · · focus · HN ↗
We were all teenagers having fun though. The friendships you speak of never really took place for me. I wish they did, because it was always jarring to unplug and get back to the real world (school, uni, work) and associate with people who had no idea of my other parallel life online.
GodelNumbering · · focus · HN ↗
[1] <a href="https://rolandgao.com/blog/gobench/" rel="nofollow">https://rolandgao.com/blog/gobench/
nullc · · focus · HN ↗
Programming an engine OTOH is a skill that is more general and they should all have.
Might be useful to have the target of the engine be some specific virtual machine that gets a strict cycle budget-- e.g. execution runs so many cycles, and result is read out of a specific memory address at the end (or when it terminates early).
ForHackernews · · focus · HN ↗
nullc · · focus · HN ↗
And it's also just bencmaxxing bait: you can get a huge improvement on the task by RLing on it, but make no improvement on anything else. Doing so would just waste model capacity.
If you could tell that every LLM was equally not being exposed to the task then you could justify it as a test of abstract reasoning, but you can't. So it ends up on how much go transcripts ended up in the training, which is ... not a very interesting metric.
theendisney · · focus · HN ↗
monster_truck · · focus · HN ↗
Shorel · · focus · HN ↗
(I have not reached single digit kyu)
monster_truck · · focus · HN ↗
winwang · · focus · HN ↗
benswerd · · focus · HN ↗
winwang · · focus · HN ↗
Although, maybe benchmarking an agent on "how well can you command a swarm to annihilate the Terrans" is how it all starts going downhill...
dschuessler · · focus · HN ↗
benswerd · · focus · HN ↗
orbital-decay · · focus · HN ↗
benswerd · · focus · HN ↗
loeg · · focus · HN ↗
benswerd · · focus · HN ↗
Marine staggering for example seems like an ideal code mode task.
loeg · · focus · HN ↗
jayd16 · · focus · HN ↗
swiftcoder · · focus · HN ↗
On the other hand, I do think LLM-based bots will quickly outperform the decision-making of many of the hand-coded bots, so maybe they won't need so much APM to be competitive.
jamiequint · · focus · HN ↗
bgandrew1 · · focus · HN ↗
adsfgoinoi · · focus · HN ↗
[dead]
qerghnui · · focus · HN ↗
AlphaStar won a showmatch against TLO. TLO was never one of the strongest players in the world. He had been retired for over three years by the time of the match. Google set the rule that their system would have human-like mechanics, but it played several times faster than any human, never issued a wasted action, had an inhumanly fast reaction time, issued commands with perfect accuracy using an API, and could see the entire map at once.
It was later released to the open ladder with more human-level mechanics. Even strong amateurs regularly beat it. I have beaten it myself. It was strong, but not even close to the level of the strongest human players. It had obvious and easily-exploitable deficiencies in strategy and building placement.
I think even the cheater version would have lost handily to Serral or any of the strongest players.
(It apparently beat MaNa as well as TLO, but those matches were never released to my knowledge. I see no reason to assume Google cheated less flagrantly in private than they did in public.)
gadtfly · · focus · HN ↗
Did it play by looking at screenshots and sending clicks, or was there other mediation/symbolization?
I have recently seen other harnesses letting agents play games in what seems like discrete time slices, turning eg Portal into something turn-based.
loeg · · focus · HN ↗
callmekit · · focus · HN ↗
faeyanpiraat · · focus · HN ↗
ericpruitt · · focus · HN ↗
vitaflo · · focus · HN ↗
I've watched several replays on this. It using map hacks is lame but it also doesn't always take advantage of them. It sometimes does respect its own fog of war. It mostly wins with incredible micro (kinda has to, it's macro kinda sucks).
Most of its few losses come from drops (it never makes turrets and doesn't know how to handle them), bad macro (blocking its own ramp) or just incredibly unconventional play from its opponent (which is certainly not in its training set).
chrishare · · focus · HN ↗
stephbook · · focus · HN ↗
"Oh it's able to issue commands too fast." "Oh no, I give it full map access and it uses that."
Ego shooters are, naturally, also not good benchmarks.
OpenAI won DotA2 in 2019, a way better game.
monster_truck · · focus · HN ↗
Forgeties79 · · focus · HN ↗
I don’t applaud people using aim assist in CS for the same reason. We literally don’t have access to the same tools.
stephbook · · focus · HN ↗
That's exactly why it's a bad game for human/AI competition.
iammjm · · focus · HN ↗
y-curious · · focus · HN ↗
cindyllm · · focus · HN ↗
[dead]
readams · · focus · HN ↗
monster_truck · · focus · HN ↗
tweakimp · · focus · HN ↗
minimal_action · · focus · HN ↗
WillMorr · · focus · HN ↗
Must recently I built out a MMORPG puzzle box thing, I wrote a general game architecture doc but left the specific puzzle design up to Fable. Nobody is actively playing rn but I left it up at bot.willmorrison.net.
AntiRush · · focus · HN ↗
<a href="https://web.archive.org/web/20091124210529/http://eis.ucsc.edu/StarCraftAICompetition" rel="nofollow">https://web.archive.org/web/20091124210529/http://eis.ucsc.e...
There's a great contemporary Ars Technica piece by a competitor:
<a href="https://arstechnica.com/gaming/2011/01/skynet-meets-the-swarm-how-the-berkeley-overmind-won-the-2010-starcraft-ai-competition/" rel="nofollow">https://arstechnica.com/gaming/2011/01/skynet-meets-the-swar...
As an undergrad I did a project using genetic programming. It was not very successful, but it was a lot of fun.
<a href="https://tomisin.space/archive/starcraft-genetic-programming/" rel="nofollow">https://tomisin.space/archive/starcraft-genetic-programming/
dcl · · focus · HN ↗
tinco · · focus · HN ↗
Unfortunately the proxybot limitations precluded me from expanding its capabilities so I tried to switch to having ruby embedded in C++ and I basically got mired there and was eventually distracted by real world concerns like actually finishing my degree.
I think I played against Krasi0's bot a couple times in the early days. Hopefully they stuck with AI and have suddenly become crazy rich after 2016. It certainly wasn't a given that AI was going to lead to a fruitful career back then, let alone to unimaginable riches.
mslate · · focus · HN ↗
<a href="https://theaccidentalengineer.com/adversarial-machine-learning-ben-weber-zynga/" rel="nofollow">https://theaccidentalengineer.com/adversarial-machine-learni...
c7b · · focus · HN ↗
chaostheory · · focus · HN ↗
mococa · · focus · HN ↗
moomin · · focus · HN ↗
therealdrag0 · · focus · HN ↗
aetherspawn · · focus · HN ↗
mcteamster · · focus · HN ↗
Protoss: powerful and expensive frontier coding agents you directly micromanage for the toughest tasks
Terran: versatile team comps of dedicated agent roles you can delegate well-defined tasks to
Zerg: massive swarms of specialist custom agents inside your apps that you evolve and optimise for speed and cost
Knowing every faction has its strengths and weaknesses helps me decide which tools to use for the job.
alchemism · · focus · HN ↗
mcteamster · · focus · HN ↗
alchemism · · focus · HN ↗
Barrin92 · · focus · HN ↗
I saw someone recently try to get an agentic system to play Final Fantasy and it did about as well as a Roomba.
ethanpailes · · focus · HN ↗
If you just mean general LLMs can’t beat humans yet that’s one thing, but it’s not the case that no system can do so.
Barrin92 · · focus · HN ↗
suby · · focus · HN ↗
silentkat · · focus · HN ↗
sharkjacobs · · focus · HN ↗
efxhoy · · focus · HN ↗
KeplerBoy · · focus · HN ↗
3eb7988a1663 · · focus · HN ↗
duskwuff · · focus · HN ↗
3eb7988a1663 · · focus · HN ↗
I was wondering how you could possibly sync up something like a Carrier that has multiple fighters.
duskwuff · · focus · HN ↗
3eb7988a1663 · · focus · HN ↗
duskwuff · · focus · HN ↗
navane · · focus · HN ↗
nemo1618 · · focus · HN ↗
mountainplus · · focus · HN ↗
The key expression for playlists or yt browsing is this `AI 업스케일`
f.e. <a href="https://www.youtube.com/playlist?list=PLWQeRMoEALvp75uolIpIEEuwsGHsiw0cu" rel="nofollow">https://www.youtube.com/playlist?list=PLWQeRMoEALvp75uolIpIE...
DJMolehill · · focus · HN ↗
windowshopping · · focus · HN ↗
xyzsparetimexyz · · focus · HN ↗
American87 · · focus · HN ↗
Svenstaro · · focus · HN ↗
snikeris · · focus · HN ↗
[deleted] · · focus · HN ↗
[deleted]
bigcat12345678 · · focus · HN ↗
kevinrineer · · focus · HN ↗
rubiquity · · focus · HN ↗
karim79 · · focus · HN ↗
I will always love this and now I'm going to play it again. Remastered and on a fancy modern machine.
iririririr · · focus · HN ↗
karim79 · · focus · HN ↗
Brood War was the first video game to be broadcast on TV in Korea. I'm pretty sure it's still going.
vanderZwan · · focus · HN ↗
stymaar · · focus · HN ↗
mkotlikov · · focus · HN ↗
ashdnazg · · focus · HN ↗
sqrt_1 · · focus · HN ↗
Narishma · · focus · HN ↗
agentdev001 · · focus · HN ↗
Karliss · · focus · HN ↗
Havoc · · focus · HN ↗
Surprised the outcomes are so poor though. I recall years ago AI was capable of beating pro level DOTA teams.
I guess in one case it was specifically trained on the interface & game while here it was not?
mrkeen · · focus · HN ↗
The film Shrek 3 (2007) took 20 million CPU hours of render time. Games push 60 frames a second. For a visual comparison, check Call of Duty world at war (2008).
Also compare to browsers, which can sometimes scroll smoothly through some styled rectangles and text, and consume gigabytes of ram if you have a few tabs open.
Games have directional sound effects and soundtracks. Don't need 800 Spotify engineers to pull that off.
Multiplayer games solve crazy distributed system problems, making it feel like 'now' when players shoot each other, even with historical latencies of 100-200ms.
AI (in terms of LLMs) seems to be a continuation of that. You used to be able to play 7 AIs on 1998 hardware, at a distinctly "non-beginner level".
TeMPOraL · · focus · HN ↗
That barrel you shoot, is really half a barrel when you're up close, a flat rectangle when you're far, a point-with-mass + a vector for purposes of physics, and not even there for purposes of AI because pathfinding uses a precomputed graph of nodes that's carefully aligned with the map so you don't notice the enemies can noclip through everything other than floors and walls. Etc.
I grew up wanting to make games, spent my teenage years in hobbyist gamedev communities, and to date, this remains to me the most enjoyable and pure form of exercising software development skills.
cm2012 · · focus · HN ↗
alembic_fumes · · focus · HN ↗
I'm often asking myself is it better to use higher or lower effort levels, or to maybe drop down to a "dumber" but faster model. And so using a real-time based competition as a benchmark could shed some light on this, I think.
In this vein, here are what I would love to see added in this benchmark:
- Include Google's Gemini models. I keep hearing Gemini being praised for its speed, and I would like to see whether that gives it a big enough edge over the bigger but slower models. - How does a Cerebras-accelerated open source model fare against a much larger but much slower frontier model?
I also feel like in general there is a lot of very low-hanging fruit to start benchmarking models across the spectrum of real-time vs batch-style workloads. Perhaps Brood War sits somewhere quite near the "real-time" end of the spectrum, but what about something like a game of speed chess, or a turn-based game with time limits?
I think what I would like to see the most is for someone to come up with a benchmark that supports tuning the "real-timeliness" of the benchmark, and then running a sweep of a model across the whole spectrum. That could get result in real nice graphs with multiple models on the pareto-frontier, varying based on the hosting provider and the model dimensions.
tianqi · · focus · HN ↗
That’s me. I’ve always struggled with real-time games because I need to pause and think. While I excel at chess and board games, I’m just no good at real-time ones. At last I can only manage by sticking to a fixed set of tactics for a game, which minimizes the need for on-the-fly thinking. Seeing current models face the same difficulty leaves me with mixed feelings.
27388383 · · focus · HN ↗
aswegs8 · · focus · HN ↗
herodoturtle · · focus · HN ↗
egeozcan · · focus · HN ↗
kennywinker · · focus · HN ↗
Like if a model can’t handle a task it hasn’t been trained on extensively, that’s not intelligence it’s memorization.
egeozcan · · focus · HN ↗
steve_taylor · · focus · HN ↗
DeepYogurt · · focus · HN ↗
rob313 · · focus · HN ↗
Really looking forward to playing- are you all limiting boxes?
leobuskin · · focus · HN ↗
I’ve scrolled the article, but haven’t noticed any remarks about Fable’s levels.
monk_grilla · · focus · HN ↗
steve_taylor · · focus · HN ↗
nrightnour · · focus · HN ↗