FYI VSCode's SSH Agent is a godsend for remote development - the "disadvantages" that Fly lists are part of its advantages. I've worked in several teams that have made extensive use of the extension, and it's never been an issue. You can restrict SSH access arbitrarily to ensure whatever security or access guardrails you need.
As a Linux user I've hated VSCode's ssh. There's lot of annoying things that make it harder to admin for. Like it doesn't pick up the MotD, preventing me from showing users important messages. I've found that it also doesn't reuse sessions (at least by default. TBF, neither does ssh) and I'll find that there's just dozens of open sessions over months from users. I literally had to write a script to boot people...
It would be one thing if the plugin was just a wrapper and people were still expected to know ssh but the plugin abstracts away all that and is intended to make it a "use VSCode on remote machine" tool. So it needs to do more than just handle creds, otherwise it creates a divergent experience while making people think it's just ssh
Exactly! Like if I'm admining a server what am I supposed to do? Message on a big slack channel and have everyone ignore me? It's easy when people are just logging in through normal ssh as I can put a big bright warning message on their screen that they can't ignore.
I have looked for mosh support for a while and not found anything. It would drastically improve the connection experience in VSCode. My terminals never disconnect anymore, but the Code popups about your sessions needing to be restarted has drastically reduced my usage.
But so is the question. MOSH interprets all the escape sequences and uses them to decide what to send to the client. That way if you tail a log file and then get disconnected, you don’t have to download every line of text that was output to your terminal while you were away; it can just send you what is currently visible.
MOSH is strictly for interactive use; never ever for automated uses like TRAMP or VSCode or sshfs.
I don't find it that opaque. Even without trying to deobfuscate the obfuscated source code which Microsoft ships (I haven't tried but it wouldn't be hard) a lot of details about how it works become obvious just by reading its logs.
Of course, it is a pity Microsoft doesn't open source it. But there are some well-maintained open source alternatives, e.g. <a href="https://github.com/jeanp413/open-remote-ssh" rel="nofollow">https://github.com/jeanp413/open-remote-ssh and <a href="https://github.com/F1yingWhite/fast-remote-ssh" rel="nofollow">https://github.com/F1yingWhite/fast-remote-ssh (I haven't got around to giving either of them a go–but I really should.)
They will never open source it. For the same reason Pylance etc. aren't open source and MS tries hard to prevent them to be used in VSCodium. Every one of their open sourced projects contains a closed source plug that MS can pull at any time that is one of the features that gives the project its unique selling points.
> There's lot of annoying things that make it harder to admin for.
It also (AIUI) tries to walk the entire file tree, so have fun with NFS (auto)mounts.
It also amounts to letting off fork bombs: we set up limits for a maximum of 256 process per UID, and regularly get folks asking "what does this 'cannot fork' message mean?": it mean you're trying to DoS the system.
> we set up limits for a maximum of 256 process per UID, and regularly get folks asking "what does this 'cannot fork' message mean?"
The max limit on 64 bit systems is what, 4,194,303? So if you have over 16,000 users per VM this limit makes sense, otherwise it just seems user-hostile.
Every process takes some memory and other resource, yes a stale process will pretty much all end up all paged out and not massively in the way of active processes, but they still aren't entirely free so it is more than a bean-counting number.
Yes, under Linux (and most unix-a-like systems) small processes are cheap to bring up and tear down which is why we create them so much, and it is not uncommon for complex interactive commands and bits of shell scripts to create several¹, but these are all likely to be short-lived so a limit of 256 certainly doesn't seem to be obscenely low to me.
What could it be doing that requires 256+ processes to be kept around for a prolonged time?
--------
[1] made up example: comparing filtered content of two gzipped files and sending the result through a script to send alerts by mail if certain things are found would be 7+ (2x gzip, 2x or more grep, diff, bash, mail or curl depending on what service you are sending alerts through)
> The max limit on 64 bit systems is what, 4,194,303? So if you have over 16,000 users per VM this limit makes sense, otherwise it just seems user-hostile.
And yet we still regularly loads of >100 on our 64 core HPC login codes, and swap is regularly used even with 96G of system memory (we have per UID memory limits too).
What's hostile is the VSCode (and Codex and Claude) makers developing tools that basically DoS a system because they assume it will operate only on single-user machines.
(And WTF are you doing that you're forking 256 processes? We have quite a few expensive HPC nodes: use those to build, not the damn login nodes.)
> The max limit on 64 bit systems is what, 4,194,303?
What a weird framing... I'm not sure what you're even trying to argue. I mean a single process can overload the machine. Just because you can label 4m processes doesn't mean you can actually run that many programs. Just think about that for a minute. 256 processes is pretty generous
What would truly be user hostile would be to allow so many processes per UID that a small handful (maybe using VSCode) make the system slow or unresponsive to everyone else.
Honestly every developer needs to increase those default limits, they are too low for modern development... So you are just crippling them and a proof of that is they keep getting this error while in their regular workflow
I don't even hit 256 on my desktop with a shitton of things open and 75 firefox sandbox processes, much less on a remote server. What in heaven's name are you doing to cross 256?
Honestly I think this thread has just devolved to HPC admins vs people who don't understand how shared multiuser sytems work cause they've been stuck on a laptop for too long to remember.
> Honestly every developer needs to increase those default limits, they are too low for modern development...
The users can develop on >100 compute nodes, but choose not to bother doing any kind of forwarding/proxying/jumping to them and just do stuff on the login nodes.
If they can't be bothered to do a "ssh -J …" then it's on them. The resources are there.
Ah, I see. Our admins typically understood that users were being given disk quotas precisely so that they could use that disk space, but that's probably not a universal stance.
Quotas mean "more than this is clearly too much" not "please use this space".
In good times nobody minds, but in bad times when you just can't extend the drive you have to tell people off. Part of the job of making sure the system can continue to work for what you need it, under _real_ constraints.
You may have to delete stuff, you may have to shut the server down to save power. You may have to limit clock speeds. It depends on the environment and "it really should work because it should be covered by next day on site warranty and you could download more ram" often doesn't apply.
Back in the day, because the architecture was shared memory and cooperative multitasking, this is how MacOS applications declared their memory constraints.
Apps had a (recommended and then user-configurable) "Minimum memory" and "Preferred memory." The app would not launch if the OS couldn't give it the minimum. It would then give the app up to the preferred amount, if available, exclusively... This was in the era before virtual memory and paging, so there was no easy way to share memory across an application boundary.
This mean that savvy users with high-resource tasks knew you had to launch your apps in a certain order to get the architecture into the configuration to do their work.
> Quotas mean "more than this is clearly too much" not "please use this space".
Ah, memories of Uni, back when storage was fairly expensive, where we had both hard and soft quotas. The soft quota would allow for temporary growth of build artefacts and things¹ but you would get stern emails if you were over your soft quota for 24 hours, and if you persisted without good reason³ your hard quota would be reduced so you effectively have no soft quota any more.
--------
[1] some machines had no local storage that the user could touch so putting them there was not always possible, some people on Windows machines had local storage but didn't have the relevant tools locally so were actually running things on the shared server(s)² instead of that just being a storage resource
[2] via telnet/rsh/rlogin: yes, I am that old… SSH was a thing by that point, though OpenSSH wasn't, and I was using it where available, but the use of older plain-text protocols was still far far more common
[3] it wasn't actually difficult to justify a quota extension for project work, in fact people enrolled on certain modules got higher quotas automatically
Yes, let's hope it doesn't get even more expensive. Around here we are purposely not upgrading hardware unless we have to, hoping to ride through the price increases. I'm sure lots of people are doing the same thing, which probably won't help much once prices start to come back down... assuming they do.
This is a fascinatingly "pets" approach to admin; it's been ages since I've been somewhere that used this approach. I've been in the "cattle fields" for decades now.
In my ecosystem, developers don't have time to glad-hand like this; if there are issues with DEV_NODE_CFG_1_29875, we might talk about it over Slack (because I'm the first one to know there's a problem as the end-user), and if we can't sort it out they'll give me some time to backup, blank the whole machine, and I get a brand-new image of DEV_NODE_CFG_1_29875.
I can't even tell you off the top of my head what territory the physical machine is running in or whether it's the only dev-node on that hardware.
(Broadly speaking, I think Microsoft is assuming cattle ecosystems; they're not openly-hostile to pets per se, but they have a strict ranking of the priorities because the "cattle ranches," as it were, bring in more money).
I had it seen on servers that were typically reserved for the team but there was no official booking system for those machines. When you start using the machine you would typically put some note to make sure somebody else does not overrun your long-running tests or performance measurements.
> but that's probably not a universal stance.
Sometimes you aren't really "the owner". For one example I was admining my group's server in grad school. I wanted to add quotas (we already had zfs) because people were abusing home directories but my advisor and a few members were very against it because I was "over complicating things". Their worry about me wasting time (20 minutes for all machines...) resulted in hours of yelling at people over the years. All because, surprise, a small number of people can't follow instructions and abuse systems, ruining it for everyone.
Another frequent problem we had was people using the systems while others were. They wouldn't check the machine's status. And of course you can guess that I wasn't allowed to add a scheduler.
A lot of groups do things in janky ways. Often not because they don't know any better but because leadership doesn't and is assertive
We have the servers role, you can derive that from the name obviously, but we do have hosts which as the same naming scheme, but slightly different roles. There's when the last Puppet run happened and what it applied (and who authored it). Depending on the host type there's also active/standby, warnings for production hosts or information about increased log level on things like sudo.
It sounds like a lot, but it's fairly compact and really helps when you need to absolutely sure where you are and you have eight terminal windows open.
"The announcement was visible in a MOTD on every server in the fra3 location for the last two days."
"I have not touched any machines in fra3 for a week, how the hell would I know of it? Why do we even have #fra3-maintenance and #maintenance channels then, if that's your stance?"
> "I have not touched any machines in fra3 for a week, how the hell would I know of it? Why do we even have #fra3-maintenance and #maintenance channels then, if that's your stance?"
We put the MOTD up ≥7 days in advance and put it in relevant Slack channels.
I miss fingering, it was such an easy way to get a log dump or status update from various daemons, I still think it has immense utility .. some of my fondest operator days were sat under the umbrella with a terminal while sleep 30 ; finger someone@all-the-things ; done .. watching the machines from afar.
Can still do it these days of course, but one with a seriously copious helping of ssh in the mix too ..
Trouble is, nobody else can do it. The only reason I have to use {social-media-blob} is because my friends don't know how to finger.
Things that should probably be an email. I think one time I’ve seen it work is to tell you stuff specifically about the host you’re on to avoid mistakes.
> I'll find that there's just dozens of open sessions over months from users
We've had the same issue with our local HPC; a few login nodes serving hundreds of users at a time, and each login node used to get swamped by these dangling SSH sessions/servers. They also wrote a script that shuts down all sessions once a day to save the login nodes.
The one this that is better than just SSH+Tmux+Vim, is that if your latency is higher than 30-50ms, since VSCode's SSH agent streams the files to your computer, the typing experience feels snappier. When you work half a continent away from where the servers are, it makes life nicer.
Use a local vim, and use its ssh support. It will download the file to the local buffer then upload it when you save. This way your vim config remains on your local machine, too.
Related, netrw (vim) can do quite a lot of things that I think people don't realize. I mean not just that it supports file trees, but well... just open :h netrw
It's just another messaging. I used them to automate messages about current disk usage and warn if the machine was actively being used.
But then again, I had people who would run jobs without checking if the machine is already in use. Obviously these people didn't check email or slack either...
danielklnstein · · focus · HN ↗
FYI VSCode's SSH Agent is a godsend for remote development - the "disadvantages" that Fly lists are part of its advantages. I've worked in several teams that have made extensive use of the extension, and it's never been an issue. You can restrict SSH access arbitrarily to ensure whatever security or access guardrails you need.
godelski · · focus · HN ↗
It would be one thing if the plugin was just a wrapper and people were still expected to know ssh but the plugin abstracts away all that and is intended to make it a "use VSCode on remote machine" tool. So it needs to do more than just handle creds, otherwise it creates a divergent experience while making people think it's just ssh
causal · · focus · HN ↗
godelski · · focus · HN ↗
Also, timeouts...
Also, does anyone know if VSCode supports mosh?
serbuvlad · · focus · HN ↗
godelski · · focus · HN ↗
serbuvlad · · focus · HN ↗
lexicality · · focus · HN ↗
serbuvlad · · focus · HN ↗
godelski · · focus · HN ↗
bandie91 · · focus · HN ↗
boldlybold · · focus · HN ↗
db48x · · focus · HN ↗
bogantech · · focus · HN ↗
db48x · · focus · HN ↗
But so is the question. MOSH interprets all the escape sequences and uses them to decide what to send to the client. That way if you tail a log file and then get disconnected, you don’t have to download every line of text that was output to your terminal while you were away; it can just send you what is currently visible.
MOSH is strictly for interactive use; never ever for automated uses like TRAMP or VSCode or sshfs.
skissane · · focus · HN ↗
Of course, it is a pity Microsoft doesn't open source it. But there are some well-maintained open source alternatives, e.g. <a href="https://github.com/jeanp413/open-remote-ssh" rel="nofollow">https://github.com/jeanp413/open-remote-ssh and <a href="https://github.com/F1yingWhite/fast-remote-ssh" rel="nofollow">https://github.com/F1yingWhite/fast-remote-ssh (I haven't got around to giving either of them a go–but I really should.)
AnonymousPlanet · · focus · HN ↗
throw0101a · · focus · HN ↗
It also (AIUI) tries to walk the entire file tree, so have fun with NFS (auto)mounts.
It also amounts to letting off fork bombs: we set up limits for a maximum of 256 process per UID, and regularly get folks asking "what does this 'cannot fork' message mean?": it mean you're trying to DoS the system.
astrange · · focus · HN ↗
DougBTX · · focus · HN ↗
The max limit on 64 bit systems is what, 4,194,303? So if you have over 16,000 users per VM this limit makes sense, otherwise it just seems user-hostile.
dspillett · · focus · HN ↗
Yes, under Linux (and most unix-a-like systems) small processes are cheap to bring up and tear down which is why we create them so much, and it is not uncommon for complex interactive commands and bits of shell scripts to create several¹, but these are all likely to be short-lived so a limit of 256 certainly doesn't seem to be obscenely low to me.
What could it be doing that requires 256+ processes to be kept around for a prolonged time?
--------
[1] made up example: comparing filtered content of two gzipped files and sending the result through a script to send alerts by mail if certain things are found would be 7+ (2x gzip, 2x or more grep, diff, bash, mail or curl depending on what service you are sending alerts through)
throw0101a · · focus · HN ↗
And yet we still regularly loads of >100 on our 64 core HPC login codes, and swap is regularly used even with 96G of system memory (we have per UID memory limits too).
What's hostile is the VSCode (and Codex and Claude) makers developing tools that basically DoS a system because they assume it will operate only on single-user machines.
(And WTF are you doing that you're forking 256 processes? We have quite a few expensive HPC nodes: use those to build, not the damn login nodes.)
godelski · · focus · HN ↗
pinkgolem · · focus · HN ↗
if you have a paid tier which offers more, you do you
if this is internally and you are a service provider to people.. why?
also 256 is not much today, my mac with a few things open is at 800
californical · · focus · HN ↗
Compared to someone on an ssh connection. No desktop, no Apple Account, no user session programs, etc. You really don’t need much
pinkgolem · · focus · HN ↗
i do not know why/in which context you are running this, and how frequently your users are executing forkbombs(i assume school/kids?)
my server is also running 400 something processes, one postgres instance alone is like 40?
godelski · · focus · HN ↗
bitfilped · · focus · HN ↗
cmiles74 · · focus · HN ↗
drowsspa · · focus · HN ↗
eqvinox · · focus · HN ↗
(Also this isn't a default limit.)
godelski · · focus · HN ↗
Too low? 256 reads as *pretty* generous to me.
bitfilped · · focus · HN ↗
throw0101a · · focus · HN ↗
The users can develop on >100 compute nodes, but choose not to bother doing any kind of forwarding/proxying/jumping to them and just do stuff on the login nodes.
If they can't be bothered to do a "ssh -J …" then it's on them. The resources are there.
Joker_vD · · focus · HN ↗
godelski · · focus · HN ↗
Joker_vD · · focus · HN ↗
literalAardvark · · focus · HN ↗
Quotas mean "more than this is clearly too much" not "please use this space".
In good times nobody minds, but in bad times when you just can't extend the drive you have to tell people off. Part of the job of making sure the system can continue to work for what you need it, under _real_ constraints.
You may have to delete stuff, you may have to shut the server down to save power. You may have to limit clock speeds. It depends on the environment and "it really should work because it should be covered by next day on site warranty and you could download more ram" often doesn't apply.
williamdclt · · focus · HN ↗
super tangential, but makes me think I never realised that "quota" can either be a lower or an upper bound depending on context
shadowgovt · · focus · HN ↗
Apps had a (recommended and then user-configurable) "Minimum memory" and "Preferred memory." The app would not launch if the OS couldn't give it the minimum. It would then give the app up to the preferred amount, if available, exclusively... This was in the era before virtual memory and paging, so there was no easy way to share memory across an application boundary.
This mean that savvy users with high-resource tasks knew you had to launch your apps in a certain order to get the architecture into the configuration to do their work.
dspillett · · focus · HN ↗
Ah, memories of Uni, back when storage was fairly expensive, where we had both hard and soft quotas. The soft quota would allow for temporary growth of build artefacts and things¹ but you would get stern emails if you were over your soft quota for 24 hours, and if you persisted without good reason³ your hard quota would be reduced so you effectively have no soft quota any more.
--------
[1] some machines had no local storage that the user could touch so putting them there was not always possible, some people on Windows machines had local storage but didn't have the relevant tools locally so were actually running things on the shared server(s)² instead of that just being a storage resource
[2] via telnet/rsh/rlogin: yes, I am that old… SSH was a thing by that point, though OpenSSH wasn't, and I was using it where available, but the use of older plain-text protocols was still far far more common
[3] it wasn't actually difficult to justify a quota extension for project work, in fact people enrolled on certain modules got higher quotas automatically
zie · · focus · HN ↗
Don't worry, disk got expensive again. Disks are usually at least 2X more expensive than a year ago currently. Sometimes 3X more.
dspillett · · focus · HN ↗
zie · · focus · HN ↗
wongarsu · · focus · HN ↗
HDD prices are a bit more reasonable. Mostly because they are so heavy. Still insane to where prices were just 12 months ago
throw0101a · · focus · HN ↗
Probably why both soft and hard limits were developed.
shadowgovt · · focus · HN ↗
In my ecosystem, developers don't have time to glad-hand like this; if there are issues with DEV_NODE_CFG_1_29875, we might talk about it over Slack (because I'm the first one to know there's a problem as the end-user), and if we can't sort it out they'll give me some time to backup, blank the whole machine, and I get a brand-new image of DEV_NODE_CFG_1_29875.
I can't even tell you off the top of my head what territory the physical machine is running in or whether it's the only dev-node on that hardware.
(Broadly speaking, I think Microsoft is assuming cattle ecosystems; they're not openly-hostile to pets per se, but they have a strict ranking of the priorities because the "cattle ranches," as it were, bring in more money).
menaerus · · focus · HN ↗
godelski · · focus · HN ↗
Another frequent problem we had was people using the systems while others were. They wouldn't check the machine's status. And of course you can guess that I wasn't allowed to add a scheduler.
A lot of groups do things in janky ways. Often not because they don't know any better but because leadership doesn't and is assertive
mrweasel · · focus · HN ↗
It sounds like a lot, but it's fairly compact and really helps when you need to absolutely sure where you are and you have eight terminal windows open.
throw0101a · · focus · HN ↗
Joker_vD · · focus · HN ↗
"I have not touched any machines in fra3 for a week, how the hell would I know of it? Why do we even have #fra3-maintenance and #maintenance channels then, if that's your stance?"
godelski · · focus · HN ↗
throw0101a · · focus · HN ↗
We put the MOTD up ≥7 days in advance and put it in relevant Slack channels.
We still get "Is Foo down?" the day of.
Quarrel · · focus · HN ↗
Like, I used to, but it was in the early 1990s.. People look at them now?
Next I'll be asking people to finger me to get my availability ..
MomsAVoxell · · focus · HN ↗
Can still do it these days of course, but one with a seriously copious helping of ssh in the mix too ..
Trouble is, nobody else can do it. The only reason I have to use {social-media-blob} is because my friends don't know how to finger.
Waterluvian · · focus · HN ↗
shadowgovt · · focus · HN ↗
godelski · · focus · HN ↗
a_bonobo · · focus · HN ↗
We've had the same issue with our local HPC; a few login nodes serving hundreds of users at a time, and each login node used to get swamped by these dangling SSH sessions/servers. They also wrote a script that shuts down all sessions once a day to save the login nodes.
w4der · · focus · HN ↗
sneak · · focus · HN ↗
godelski · · focus · HN ↗
w4der · · focus · HN ↗
user43928 · · focus · HN ↗
shadowgovt · · focus · HN ↗
No snark: does your org also regularly check the mail spool and expect individual users to do so as well?
godelski · · focus · HN ↗
But then again, I had people who would run jobs without checking if the machine is already in use. Obviously these people didn't check email or slack either...