I strongly believe A/B testing users without their knowledge and enthusiastic consent is unethical. If you don't have enough confidence in your changes to make them carte blanche, then don't make them at all. Users don't want their shit changing all the time.
> A/B testing users without their knowledge and enthusiastic consent is unethical.
Yep, and A/B testing as experienced by uninformed, unaware end-users is a dark pattern.
It undermines the perception of (and trust in) continuity which is necessary to make effective use of a tool. The best way I can describe it to the skeptical is: imagine the dials on your car's dashboard rearrange themselves occasionally overnight, and on some commutes to work you suddenly can't work the radio or the AC while moving at ≥35mph. Of course, since the widespread use of touchscreens, that example became very literal.
So the car manufacturer has figured out the "optimal" arrangement of dials and buttons on their dashboard for their preferred levels of user engagement. Great. How many of those users now associate their car's brand with inconsistency? "I can't trust the damn buttons to be in the same place the next time I drive."
I think this is a bad analogy (or a good analogy for bad A/B testing).
I would say it's more like: imagine if your dashboard controls and icons, had occasional tweaks in size and shape that made them slightly harder/easier to use, but over time ended up with controls you found more intuitive and easier to use.
> ...imagine if your dashboard controls and icons, had occasional tweaks in size and shape that made them slightly harder/easier to use, but over time ended up with controls you found more intuitive and easier to use.
When even a single button disappears from where I expect to find it in an application (and reappears somewhere else), that has never resulted in the application feeling more intuitive or easy to use. It has only ever been frustrating in the most literal sense of that word, and only every caused me to start looking into other programs.
A more subtle undermining of the user's trust that their tools will be where they left them on the screen the next time they open the application is arguably more insidious, not less. Both because it is more maddening to the end-user than a full redesign and because the damage of lost trust and user ire is harder to measure than the A/B engagement data until things progress too far for any easy, non-painful remediation.
A/B testing can be helpful as a part of an actual user study on users who are aware of what is happening. But unleashed on an unsuspecting user population out in the real world with real things to do on real deadlines, A/B testing inevitably becomes a harmful dark pattern.
I would never want this unless I needed to give clear and obvious consent to enable it in the first place, and have access to a 'reset to default and never change again' button. Ideally I'd probably also have 'good change' and 'bad change' buttons to point it in the right direction.
This is so far removed from actual A/B testing that happens in the real world that it's unhelpful to even pretend it could be.
There are small A/B tests which absolutely make sense. You often see marketing sites making small tweaks to banners and copy. It's not that one has worse UX or even that one is objectively worse, just that different users have different preferences and it's often difficult to know exactly what will work best.
Similarly you can be very confident of something, but A/B testing it still reduces risk. Any significant change should probably always be rolled out to a small fraction of the user base first in case you accidentally change something for the worse.
I agree if you're talking about some BS experiment where a company uses A/B testing as an alternative to putting the hours into product design and user research.
I agree. I work at a big insurance company with millions of online visitors each year. By A/B testing, we improve our services. A/B testing error messages has been a huge help for us. We can't ask users directly what they need: most users don't know what they need. That's why we do both qualitative (user research, online feedback forms) and quantitative (A/B testing, fake door testing, etc.) research.
Hmmm, I am curious about which aspect of this you find unethical.
Is it unethical to do phased rollouts (where a small percentage get the new version) as a way to do safe deploys? If the issue is that two users making requests at the same time might see different things, then this would also be unethical? Yet, these sorts of phases rollouts is the best way to release something safely. When I worked at a large CDN with 50,000 servers around the world, we ALWAYS did phased releases, to make sure we didn't take down everything all at once, and to make sure we caught any performance regressions right away.
Is your issue that the user might be getting a version that won't stick around? That seems always the case, whether you do A/B or not. You might rollback if there is an issue, and you will certainly roll forward at some point, meaning users will get a new version at some point.
Would it be an issue if the A/B test was temporal? Like all users got one version today, and a different version tomorrow?
I guess I am just confused by this statement:
> If you don't have enough confidence in your changes to make them carte blanche, then don't make them at all.
This seems contrary to so many other complaints we see all the time, that companies push changes out without taking into account what users actually want. So, do we want companies that push out changes with no user feedback because they are confident that they know what users want, or do we want companies that get feedback from users on whether new changes are helping or hurting.
How about actually asking the users instead of experimenting on them?
> or do we want companies that get feedback from users on whether new changes are helping or hurting.
You don't get that feedback. The feedback you get is whether some telemetry KPI goes up or down. That's not the same as actual utility for the user.
Asking users for feedback is notoriously bad at generating good feedback. Most people don't respond, and those that do ask for things they don't actually want, or are only wanted by very few people.
I've heard we have this amazing new technology now that can understand natural language and automatically convert it into structured data if we ask it to. Maybe we could use that to ask people at scale now.
> and those that do ask for things they don't actually want, or are only wanted by very few people.
But you can’t build a business on “wants” that people don’t act on.
If everybody says they want foo and don’t want bar, but when you make foo they don’t buy any but will buy a lot of bar, then are you “failing to provide what people want” if you just make bar?
Obviously phased rollouts are fine and even necessary like you said. That’s not really “experimenting on people”, it’s experimenting on the network/system.
I constantly experiment on users in my work. It’s all around extracting the most money you possibly can. Meanwhile we have mountains of UX interview material where people tell us exactly what’s wrong with our site, and we don’t implement any of it lol.
Profits are up though! In a big way! And our users continue to hate us more and more.
> This seems contrary to so many other complaints we see all the time, that companies push changes out without taking into account what users actually want
Are you not aware how unpopular practically all recent changes on YT are among its users? Almost none of the changes done on YT in the last 5+ years would have happened if they took into account what users & creators want. So how exactly does telemetry and A/B testing help when it either tells them the opposite of reality or they simply interpret the data however they like anyway?
I am not saying you are wrong, but how are you so sure that you know what the average user and creator wants? There are millions of youtube users and creators who aren't participating in whatever forum you are basing your information on. Maybe the people you hear complaining are a vocal minority.
It could be that YT is just making everything worse for everyone, but I also know they have data that you and I don't have on how people actually use their product. I don't think we can assume they are just bad at making a product just because all the people we talk to agree with us that it is bad.
YouTube's goals are not aligned with what is best for the majority of users. They are optimizing for maximum time spent watching ads, hence they are working against user interests.
If I paid for a copy of Windows 95, I'm expecting Windows 95 in a box.
When I have two free hours to turn on the XBOX for the first time in half a year, I want it to turn on right away and play my game. I don't want the box I paid for and have been looking at to figure that it needs an OS update, and a game update, and that the game should now be slower and glitchier than it was the last time I played it.
When I play music on my phone in my car, I don't want to find out that the "Start Mix" button moved, or that showing the upcoming playlist now takes one more swipe, or that the UI won't load because YouTube Music doesn't cache it's UI anymore and when you have a cell network reporting 1 bar but it's actually zero bars, you get a spinner for ten minutes.
The six CD changer in my dashboard has worked exactly the same way since 2006, the discs play when the key turns on, and nothing moves. There's no engagement to be had other than "my music plays when I turn on the car in my driveway which also has spotty cell service". There isn't a KPI to be measured, a PM to be promoted, or anything. It's just a car radio.
Most consumer goods are solved problems. Nothing's changed since 2015. Even tech from 2015 is just a convergence of 2005 tech, like MP3 players, digital cameras, and Blackberries. People don't have radically different problems to solve in their day to day lives. People take pictures, share them, do email and group chat, voice calls, read the news, watch TV, pay for parking, do some banking. Watch a 90s TV show and all of those activites required different physical places and tactile goods. (Hell, that's why screens in Android are called *Activities*.).
When I pay for a product, I expect to be paying for a finished product, not some psychological experiment that's someone else's promo packet.
Most of the websites people are talking about here are free, they aren't things you paid for, so the argument that when you pay for a product you should expect something doesn't really apply.
I think an important distinction being missed by those relying to you is that you seem to be saying that A/B testing is unethical when it is used to optimize human resource extraction (i.e. marketing).
(I have this weird feeling that you're not upset with blue/green deployments...)
> Oh, and how much are you paying for that service...?
I paid YT Premium once. It unlocked a playback queue in the app. That queue had 5 separate bugs I found within an hour. It's literally a simple playlist and yet not a single feature (adding, removing, reordering etc) worked reliably. The playback queue also randomly emptied itself sometimes. Meanwhile I get a superior version of this in the browser by simply opening a video in another tab, for free.
Why should I pay for that while they only have AI support designed to never solve any issues, do nothing against bots, and then warn you that you may get banned when you report too many bots?
God that playback queue has been so awfully buggy for so long. Rant incoming.
I actually discussed it with a friend a few years back while debating the declining quality of Google's engineering.
The craziest thing to me - it appears to be using some kind of eventual consistency, so additions/reorders/deletions have to go through some complex process server-side that takes several seconds to update in the UI (and often the order is wrong, or silently fails to add videos). And yet, the whole queue disappears without a trace if the YouTube app gets unloaded by iOS, and is unavailable on other devices, so it could have just been stored locally all along.
I thought that perhaps they were storing it server-side because eventually it would allow the queue to be restored or transferred, but it has been about 3 years now and I don't think that day is ever coming. I just lost a whole queue of several vids this morning (on the plus side it was a good incentive to get off YT, so I'll give it that).
My guess is it's using eventual consistency or some complex multi-microservice chain of RPCs because that's just what you have to do at Google. I'm sure there are engineers who want to fix the feature or go back and complete the rushed launch but likely can't convince the decision makers that it's necessary.
So we get a subpar experience from one of the largest companies in the world with thousands of the best engineers, while Google keeps getting their $16/month because there are no alternatives
> So we get a subpar experience from one of the largest companies in the world with thousands of the best engineers, while Google keeps getting their $16/month because there are no alternatives
When it comes to paying for video streaming, they are plenty of other services, and they all provide a better user experience than Youtube. Even if I pay, say, Paramount for their ad-laden option (the cheaper one), it's superior to Youtube without ads.
In the case of youtube? I am paying quite a bit actually.
Updates are one thing, changes made specifically targeted towards manipulating users into spending more time on the site in ways they can’t opt out of (youtube shorts is a good example) are different. No one is bothered if YouTube updates to allow 8k streams. I am extremely bothered that YouTube does not allow me to disable shorts in the app. I don’t want shorts, they are a distraction and one more thing i have to guard against getting sucked into. Let me use the app how I want to use it, not how you want me to use it.
> I strongly believe A/B testing users without their knowledge and enthusiastic consent is unethical
Look I get it feels weird but in practice most A/B tests are stuff like “does this copy change if ppl use this feature”.
The reasons they don’t is the same reason RCTs for new drugs don’t tell patients either. You end up with selection bias.
> If you don't have enough confidence in your changes to make them carte blanche, then don't make them at all
This is a bit hyperbolic, empirics is something that should be used more by decision makers not just for their own sake but for people who don’t understand why they are making them, especially in government (although It’s harder because finding cases where it’s appropriate is hard).
An A/B test isn’t just about what’s better, it’s about understanding all other things being equal how does one change to X affect Y. Which is information that can be used to inform the design of yet to be build features.
A lot of ppl have bad takes on what makes a product better, and they would otherwise have a greater say in the product design. Some product managers are just really stupid and are there due to nepotism so it’s an external equaliser and allowing the thoughtful ones to have more of a say.
> Users don't want their shit changing all the time.
Yep that’s why you don’t ask them.
I get if you have a specific flow that your use to. It would annoying for me too if that changed (as I’m pretty stubborn don’t like ppl making changes on my behalf), but that doesn’t mean it’s an objective better experience for all users or users who have yet to be familiar with the apps process.
When these products operate in competitive markets and not some winner takes all market these are often about improving users experience.
If this was something more high stakes like a medial trial I’d get it, but for stuff like filling out a document or watching a piece of media. The stakes for most SASS app are really low.
> This is a bit hyperbolic, empirics is something that should be used more by decision makers
In my experience working in software, empiricism is the only thing valued anymore. Intuition and thoughtfulness is out the window because it's not scientific enough. I would say most software now reflects that - it's almost all bland and statistically optimized to maximize engagement or revenue.
> If this was something more high stakes like a medial trial I’d get it, but for stuff like filling out a document or watching a piece of media. The stakes for most SASS app are really low.
There are plenty of subtle patterns used in SaaS form filling things too. For example notice that the "primary button" is always chosen as the one that will make the company the most money or collect the most data.
Likewise, popups are annoying, but they result in more conversions. Forced logins are the same (how many form filling apps now force you to sign up with an account that you'll never use again, so that you can become a potential lead in future).
Google recently started doing all of these on anonymous searches with a modal overlay and a big bright blue "Continue" primary button that takes you to a login screen, while a "don't sign in" button appears as far less noticeable text above it.
It's at the point now where I'm surprised when any software gives you an option without blatantly telling you which one they want you to pick for their own benefit.
A lot of it seems innocuous but I feel we're at the stage of death-by-a-thousand-cuts at this point.
> In my experience working in software, empiricism is the only thing valued anymore. Intuition and thoughtfulness is out the window because it's not scientific enough.
With things like A/B testing, its not entirely an objective as you need to make assumptions which can be difficult to measure (although typically randomisation solves a lot of them), but you can only measure what you've decide to measure (which isn't random), so you don't know when you're in a local max. So IMO intuition and thoughtfulness is necessary. Sometimes product managers don't listen to data scientists when they say you can't measure Y with X, or the research design violates the required assumptions to make a causal claims (like reverse causality or controlling on a post treatment effect, e.g. employment as control when measuring income after hospitalisation (the treatment)). I think the worse offences I've seen have been from marketing teams.
But proper research design does require intuition and thoughtfulness, because statistical models require thought, like other forms of supervised learning.
I've seen both
- PMs use questionable experiments to justify shipping something.
- PMs dismiss experiments when it was a null result and shipped anyways.
In either case I don't think the methodology is the cause of problems here, although I think shipping with a null result is justifiable if it's a larger unit of work (provided its not a regression).
Sometimes things that have heterogenous effects get measured as a homogenous effect, Like say:
- Your primary user base is X1 and X2 is a larger consumer base but makes up a small portion of your user base.
- Your experiment does poorly with X1, but say there was an increase in user base X2.
- However because X1 dominates the user base and your signups (because say you target ads to X1 over X2), no one drills into the effects on these different user bases, the result gets discarded as it seems to be a bad outcome.
There's valuable information in the experiment outcome but without thought and attention you can miss it.
> Google recently started doing all of these on anonymous searches with a modal overlay and a big bright blue "Continue" primary button that takes you to a login screen, while a "don't sign in" button appears as far less noticeable text above it.
I mean that sucks, but IMO with their market share, the way Google chrome is inclined to develop their product is very different to firms in more competitive spaces.
Perhaps A/B Tests, allows google to optimise the things they are incentivised to pursue, but in the hands of smaller firms with different incentives are willing to tweak things to be more appealing to users when they have far less market power, which I think is probably more the issue in the case of Google.
I just don't think this is a universal problem with the methodology
Oh hey didn't check your name before I replied! I assume this is the akst I know IRL, small world. Hope you are well too.
> I just don't think this is a universal problem with the methodology
In an isolated world I would agree with you, but in the messy reality we live in I think the broader problems with the methodology are two:
1. The belief that everything can be measured. There are many intangibles (user trust, willingness to put up with bugs, "vibes") that aren't easily measurable. Yes, net promoter scores etc, but every company I can think of that uses them builds bland, buggy, largely disliked products that people use only because they have no choice.
2. Focusing on experiments and measurable outcomes creates a tendency to de-prioritise anything that isn't easily measurable. For example larger, riskier projects that can't be quickly tested. Or whimsy - easter eggs that devs added to many products in the past that users remember for years. Rarely added anymore because they can't be justified against a roadmap full of experiments.
Sometimes it feels a bit like we're sitting around so focused on measuring whether guests prefer one dish or the other and arguing over which experiment is best, that we don't notice a forest fire is blazing outside.
To me maybe it's a bit of a question of science vs art. Take the video games industry for example, the AAA studios are largely moving toward the science end of the spectrum with predictable games that extract maximum engagement, while indie studios are mostly building small, unique experiences that don't try to dominate your attention.
Experimentation always narrows focus down, and I think it's worth considering what's being traded off by doing so.
It's possible I've spent too much time looking at town planning where nothing is really ever measured and when numbers are produced it is it some insane
Like claiming townhouses not facing out into the street somehow produces X $ in mental health costs due to "lack of inclusion" and the footnote links to a study where non-english speaking communities in Australia were having real health costs due to lack of access to translation services. Which is something that passes for "evidence based" planning. Doing experiments in social sciences is a lot harder tho.
Maybe it's too easy to do experiments in Software and like you said
> Sometimes it feels a bit like we're sitting around so focused on measuring whether guests prefer one dish or the other and arguing over which experiment is best, that we don't notice a forest fire is blazing outside
I do think they're a useful tool but I see what you're saying.
Part of me feel some of that is risk aversion, but also maybe some of its process dependence when you have a number that provides strong certainty on a number of things, and because there's that feedback loop they get drawn to the things that cause the number go up. Similar to how people become dependent on LLMs to get stuff done or affirm if they did the right thing, and when they're in a space that's harder to measure they don't know to judge if they've done a good job or not.
I'm spending less time working on software, as I decided to get an econ degree, specifically econometrics, so I do spent a lot of time trying to think how to better measure stuff specifically in public policy, so I might be bias lol
But I do get what you're saying, hope things are well for you
Software changes before my eyes constantly via normal releases so I don’t really care if I’m part of an experiment or not.
That’s software though—I see your point for something like content. I’m already used to seeing the title or thumbnail of a YouTube video change as the result of an experiment “winning” but the content itself…that would very jarring.
OkayPhysicist · · focus · HN ↗
sfRattan · · focus · HN ↗
Yep, and A/B testing as experienced by uninformed, unaware end-users is a dark pattern.
It undermines the perception of (and trust in) continuity which is necessary to make effective use of a tool. The best way I can describe it to the skeptical is: imagine the dials on your car's dashboard rearrange themselves occasionally overnight, and on some commutes to work you suddenly can't work the radio or the AC while moving at ≥35mph. Of course, since the widespread use of touchscreens, that example became very literal.
So the car manufacturer has figured out the "optimal" arrangement of dials and buttons on their dashboard for their preferred levels of user engagement. Great. How many of those users now associate their car's brand with inconsistency? "I can't trust the damn buttons to be in the same place the next time I drive."
sceptic123 · · focus · HN ↗
I would say it's more like: imagine if your dashboard controls and icons, had occasional tweaks in size and shape that made them slightly harder/easier to use, but over time ended up with controls you found more intuitive and easier to use.
sfRattan · · focus · HN ↗
When even a single button disappears from where I expect to find it in an application (and reappears somewhere else), that has never resulted in the application feeling more intuitive or easy to use. It has only ever been frustrating in the most literal sense of that word, and only every caused me to start looking into other programs.
A more subtle undermining of the user's trust that their tools will be where they left them on the screen the next time they open the application is arguably more insidious, not less. Both because it is more maddening to the end-user than a full redesign and because the damage of lost trust and user ire is harder to measure than the A/B engagement data until things progress too far for any easy, non-painful remediation.
A/B testing can be helpful as a part of an actual user study on users who are aware of what is happening. But unleashed on an unsuspecting user population out in the real world with real things to do on real deadlines, A/B testing inevitably becomes a harmful dark pattern.
Telaneo · · focus · HN ↗
This is so far removed from actual A/B testing that happens in the real world that it's unhelpful to even pretend it could be.
kypro · · focus · HN ↗
There are small A/B tests which absolutely make sense. You often see marketing sites making small tweaks to banners and copy. It's not that one has worse UX or even that one is objectively worse, just that different users have different preferences and it's often difficult to know exactly what will work best.
Similarly you can be very confident of something, but A/B testing it still reduces risk. Any significant change should probably always be rolled out to a small fraction of the user base first in case you accidentally change something for the worse.
I agree if you're talking about some BS experiment where a company uses A/B testing as an alternative to putting the hours into product design and user research.
robobo96 · · focus · HN ↗
cortesoft · · focus · HN ↗
Is it unethical to do phased rollouts (where a small percentage get the new version) as a way to do safe deploys? If the issue is that two users making requests at the same time might see different things, then this would also be unethical? Yet, these sorts of phases rollouts is the best way to release something safely. When I worked at a large CDN with 50,000 servers around the world, we ALWAYS did phased releases, to make sure we didn't take down everything all at once, and to make sure we caught any performance regressions right away.
Is your issue that the user might be getting a version that won't stick around? That seems always the case, whether you do A/B or not. You might rollback if there is an issue, and you will certainly roll forward at some point, meaning users will get a new version at some point.
Would it be an issue if the A/B test was temporal? Like all users got one version today, and a different version tomorrow?
I guess I am just confused by this statement:
> If you don't have enough confidence in your changes to make them carte blanche, then don't make them at all.
This seems contrary to so many other complaints we see all the time, that companies push changes out without taking into account what users actually want. So, do we want companies that push out changes with no user feedback because they are confident that they know what users want, or do we want companies that get feedback from users on whether new changes are helping or hurting.
xg15 · · focus · HN ↗
> or do we want companies that get feedback from users on whether new changes are helping or hurting.
You don't get that feedback. The feedback you get is whether some telemetry KPI goes up or down. That's not the same as actual utility for the user.
cortesoft · · focus · HN ↗
xg15 · · focus · HN ↗
> and those that do ask for things they don't actually want, or are only wanted by very few people.
Then how do you know what people actually want?
Joker_vD · · focus · HN ↗
isabelc · · focus · HN ↗
cortesoft · · focus · HN ↗
If everybody says they want foo and don’t want bar, but when you make foo they don’t buy any but will buy a lot of bar, then are you “failing to provide what people want” if you just make bar?
fragmede · · focus · HN ↗
-(not) Henry Ford
carljungslabtek · · focus · HN ↗
I constantly experiment on users in my work. It’s all around extracting the most money you possibly can. Meanwhile we have mountains of UX interview material where people tell us exactly what’s wrong with our site, and we don’t implement any of it lol.
Profits are up though! In a big way! And our users continue to hate us more and more.
dgently7 · · focus · HN ↗
alpaca128 · · focus · HN ↗
Are you not aware how unpopular practically all recent changes on YT are among its users? Almost none of the changes done on YT in the last 5+ years would have happened if they took into account what users & creators want. So how exactly does telemetry and A/B testing help when it either tells them the opposite of reality or they simply interpret the data however they like anyway?
cortesoft · · focus · HN ↗
It could be that YT is just making everything worse for everyone, but I also know they have data that you and I don't have on how people actually use their product. I don't think we can assume they are just bad at making a product just because all the people we talk to agree with us that it is bad.
alpaca128 · · focus · HN ↗
linster · · focus · HN ↗
When I have two free hours to turn on the XBOX for the first time in half a year, I want it to turn on right away and play my game. I don't want the box I paid for and have been looking at to figure that it needs an OS update, and a game update, and that the game should now be slower and glitchier than it was the last time I played it.
When I play music on my phone in my car, I don't want to find out that the "Start Mix" button moved, or that showing the upcoming playlist now takes one more swipe, or that the UI won't load because YouTube Music doesn't cache it's UI anymore and when you have a cell network reporting 1 bar but it's actually zero bars, you get a spinner for ten minutes.
The six CD changer in my dashboard has worked exactly the same way since 2006, the discs play when the key turns on, and nothing moves. There's no engagement to be had other than "my music plays when I turn on the car in my driveway which also has spotty cell service". There isn't a KPI to be measured, a PM to be promoted, or anything. It's just a car radio.
Most consumer goods are solved problems. Nothing's changed since 2015. Even tech from 2015 is just a convergence of 2005 tech, like MP3 players, digital cameras, and Blackberries. People don't have radically different problems to solve in their day to day lives. People take pictures, share them, do email and group chat, voice calls, read the news, watch TV, pay for parking, do some banking. Watch a 90s TV show and all of those activites required different physical places and tactile goods. (Hell, that's why screens in Android are called *Activities*.).
When I pay for a product, I expect to be paying for a finished product, not some psychological experiment that's someone else's promo packet.
cortesoft · · focus · HN ↗
spaqin · · focus · HN ↗
tjpnz · · focus · HN ↗
arcanemachiner · · focus · HN ↗
(I have this weird feeling that you're not upset with blue/green deployments...)
BeetleB · · focus · HN ↗
> Users don't want their shit changing all the time.
Eliminating A/B testing won't solve this problem. Even without A/B testing, they make updates, etc.
You might as well just say "Updating an online service without asking the user first is unethical."
Oh, and how much are you paying for that service...?
pixl97 · · focus · HN ↗
alpaca128 · · focus · HN ↗
I paid YT Premium once. It unlocked a playback queue in the app. That queue had 5 separate bugs I found within an hour. It's literally a simple playlist and yet not a single feature (adding, removing, reordering etc) worked reliably. The playback queue also randomly emptied itself sometimes. Meanwhile I get a superior version of this in the browser by simply opening a video in another tab, for free.
Why should I pay for that while they only have AI support designed to never solve any issues, do nothing against bots, and then warn you that you may get banned when you report too many bots?
jordwest · · focus · HN ↗
I actually discussed it with a friend a few years back while debating the declining quality of Google's engineering.
The craziest thing to me - it appears to be using some kind of eventual consistency, so additions/reorders/deletions have to go through some complex process server-side that takes several seconds to update in the UI (and often the order is wrong, or silently fails to add videos). And yet, the whole queue disappears without a trace if the YouTube app gets unloaded by iOS, and is unavailable on other devices, so it could have just been stored locally all along.
I thought that perhaps they were storing it server-side because eventually it would allow the queue to be restored or transferred, but it has been about 3 years now and I don't think that day is ever coming. I just lost a whole queue of several vids this morning (on the plus side it was a good incentive to get off YT, so I'll give it that).
My guess is it's using eventual consistency or some complex multi-microservice chain of RPCs because that's just what you have to do at Google. I'm sure there are engineers who want to fix the feature or go back and complete the rushed launch but likely can't convince the decision makers that it's necessary.
So we get a subpar experience from one of the largest companies in the world with thousands of the best engineers, while Google keeps getting their $16/month because there are no alternatives
BeetleB · · focus · HN ↗
When it comes to paying for video streaming, they are plenty of other services, and they all provide a better user experience than Youtube. Even if I pay, say, Paramount for their ad-laden option (the cheaper one), it's superior to Youtube without ads.
Why people pay for Youtube is beyond me.
dghlsakjg · · focus · HN ↗
Updates are one thing, changes made specifically targeted towards manipulating users into spending more time on the site in ways they can’t opt out of (youtube shorts is a good example) are different. No one is bothered if YouTube updates to allow 8k streams. I am extremely bothered that YouTube does not allow me to disable shorts in the app. I don’t want shorts, they are a distraction and one more thing i have to guard against getting sucked into. Let me use the app how I want to use it, not how you want me to use it.
akst · · focus · HN ↗
Look I get it feels weird but in practice most A/B tests are stuff like “does this copy change if ppl use this feature”.
The reasons they don’t is the same reason RCTs for new drugs don’t tell patients either. You end up with selection bias.
> If you don't have enough confidence in your changes to make them carte blanche, then don't make them at all
This is a bit hyperbolic, empirics is something that should be used more by decision makers not just for their own sake but for people who don’t understand why they are making them, especially in government (although It’s harder because finding cases where it’s appropriate is hard).
An A/B test isn’t just about what’s better, it’s about understanding all other things being equal how does one change to X affect Y. Which is information that can be used to inform the design of yet to be build features.
A lot of ppl have bad takes on what makes a product better, and they would otherwise have a greater say in the product design. Some product managers are just really stupid and are there due to nepotism so it’s an external equaliser and allowing the thoughtful ones to have more of a say.
> Users don't want their shit changing all the time.
Yep that’s why you don’t ask them.
I get if you have a specific flow that your use to. It would annoying for me too if that changed (as I’m pretty stubborn don’t like ppl making changes on my behalf), but that doesn’t mean it’s an objective better experience for all users or users who have yet to be familiar with the apps process.
When these products operate in competitive markets and not some winner takes all market these are often about improving users experience.
If this was something more high stakes like a medial trial I’d get it, but for stuff like filling out a document or watching a piece of media. The stakes for most SASS app are really low.
jordwest · · focus · HN ↗
In my experience working in software, empiricism is the only thing valued anymore. Intuition and thoughtfulness is out the window because it's not scientific enough. I would say most software now reflects that - it's almost all bland and statistically optimized to maximize engagement or revenue.
> If this was something more high stakes like a medial trial I’d get it, but for stuff like filling out a document or watching a piece of media. The stakes for most SASS app are really low.
There are plenty of subtle patterns used in SaaS form filling things too. For example notice that the "primary button" is always chosen as the one that will make the company the most money or collect the most data.
Likewise, popups are annoying, but they result in more conversions. Forced logins are the same (how many form filling apps now force you to sign up with an account that you'll never use again, so that you can become a potential lead in future).
Google recently started doing all of these on anonymous searches with a modal overlay and a big bright blue "Continue" primary button that takes you to a login screen, while a "don't sign in" button appears as far less noticeable text above it.
It's at the point now where I'm surprised when any software gives you an option without blatantly telling you which one they want you to pick for their own benefit.
A lot of it seems innocuous but I feel we're at the stage of death-by-a-thousand-cuts at this point.
akst · · focus · HN ↗
> In my experience working in software, empiricism is the only thing valued anymore. Intuition and thoughtfulness is out the window because it's not scientific enough.
With things like A/B testing, its not entirely an objective as you need to make assumptions which can be difficult to measure (although typically randomisation solves a lot of them), but you can only measure what you've decide to measure (which isn't random), so you don't know when you're in a local max. So IMO intuition and thoughtfulness is necessary. Sometimes product managers don't listen to data scientists when they say you can't measure Y with X, or the research design violates the required assumptions to make a causal claims (like reverse causality or controlling on a post treatment effect, e.g. employment as control when measuring income after hospitalisation (the treatment)). I think the worse offences I've seen have been from marketing teams.
But proper research design does require intuition and thoughtfulness, because statistical models require thought, like other forms of supervised learning.
I've seen both
- PMs use questionable experiments to justify shipping something.
- PMs dismiss experiments when it was a null result and shipped anyways.
In either case I don't think the methodology is the cause of problems here, although I think shipping with a null result is justifiable if it's a larger unit of work (provided its not a regression).
Sometimes things that have heterogenous effects get measured as a homogenous effect, Like say:
- Your primary user base is X1 and X2 is a larger consumer base but makes up a small portion of your user base.
- Your experiment does poorly with X1, but say there was an increase in user base X2.
- However because X1 dominates the user base and your signups (because say you target ads to X1 over X2), no one drills into the effects on these different user bases, the result gets discarded as it seems to be a bad outcome.
There's valuable information in the experiment outcome but without thought and attention you can miss it.
> Google recently started doing all of these on anonymous searches with a modal overlay and a big bright blue "Continue" primary button that takes you to a login screen, while a "don't sign in" button appears as far less noticeable text above it.
I mean that sucks, but IMO with their market share, the way Google chrome is inclined to develop their product is very different to firms in more competitive spaces.
Perhaps A/B Tests, allows google to optimise the things they are incentivised to pursue, but in the hands of smaller firms with different incentives are willing to tweak things to be more appealing to users when they have far less market power, which I think is probably more the issue in the case of Google.
I just don't think this is a universal problem with the methodology
jordwest · · focus · HN ↗
> I just don't think this is a universal problem with the methodology
In an isolated world I would agree with you, but in the messy reality we live in I think the broader problems with the methodology are two:
1. The belief that everything can be measured. There are many intangibles (user trust, willingness to put up with bugs, "vibes") that aren't easily measurable. Yes, net promoter scores etc, but every company I can think of that uses them builds bland, buggy, largely disliked products that people use only because they have no choice.
2. Focusing on experiments and measurable outcomes creates a tendency to de-prioritise anything that isn't easily measurable. For example larger, riskier projects that can't be quickly tested. Or whimsy - easter eggs that devs added to many products in the past that users remember for years. Rarely added anymore because they can't be justified against a roadmap full of experiments.
Sometimes it feels a bit like we're sitting around so focused on measuring whether guests prefer one dish or the other and arguing over which experiment is best, that we don't notice a forest fire is blazing outside.
To me maybe it's a bit of a question of science vs art. Take the video games industry for example, the AAA studios are largely moving toward the science end of the spectrum with predictable games that extract maximum engagement, while indie studios are mostly building small, unique experiences that don't try to dominate your attention.
Experimentation always narrows focus down, and I think it's worth considering what's being traded off by doing so.
akst · · focus · HN ↗
Like claiming townhouses not facing out into the street somehow produces X $ in mental health costs due to "lack of inclusion" and the footnote links to a study where non-english speaking communities in Australia were having real health costs due to lack of access to translation services. Which is something that passes for "evidence based" planning. Doing experiments in social sciences is a lot harder tho.
Maybe it's too easy to do experiments in Software and like you said
> Sometimes it feels a bit like we're sitting around so focused on measuring whether guests prefer one dish or the other and arguing over which experiment is best, that we don't notice a forest fire is blazing outside
I do think they're a useful tool but I see what you're saying.
Part of me feel some of that is risk aversion, but also maybe some of its process dependence when you have a number that provides strong certainty on a number of things, and because there's that feedback loop they get drawn to the things that cause the number go up. Similar to how people become dependent on LLMs to get stuff done or affirm if they did the right thing, and when they're in a space that's harder to measure they don't know to judge if they've done a good job or not.
I'm spending less time working on software, as I decided to get an econ degree, specifically econometrics, so I do spent a lot of time trying to think how to better measure stuff specifically in public policy, so I might be bias lol
But I do get what you're saying, hope things are well for you
SaucyWrong · · focus · HN ↗
That’s software though—I see your point for something like content. I’m already used to seeing the title or thumbnail of a YouTube video change as the result of an experiment “winning” but the content itself…that would very jarring.