‹ BackHN Continuity

Thread

Dear Software Makers

121 points · 91 comments · speckx

  1. OkayPhysicist · · focus · HN ↗
    I strongly believe A/B testing users without their knowledge and enthusiastic consent is unethical. If you don't have enough confidence in your changes to make them carte blanche, then don't make them at all. Users don't want their shit changing all the time.
    1. akst · · focus · HN ↗
      > I strongly believe A/B testing users without their knowledge and enthusiastic consent is unethical

      Look I get it feels weird but in practice most A/B tests are stuff like “does this copy change if ppl use this feature”.

      The reasons they don’t is the same reason RCTs for new drugs don’t tell patients either. You end up with selection bias.

      > If you don't have enough confidence in your changes to make them carte blanche, then don't make them at all

      This is a bit hyperbolic, empirics is something that should be used more by decision makers not just for their own sake but for people who don’t understand why they are making them, especially in government (although It’s harder because finding cases where it’s appropriate is hard).

      An A/B test isn’t just about what’s better, it’s about understanding all other things being equal how does one change to X affect Y. Which is information that can be used to inform the design of yet to be build features.

      A lot of ppl have bad takes on what makes a product better, and they would otherwise have a greater say in the product design. Some product managers are just really stupid and are there due to nepotism so it’s an external equaliser and allowing the thoughtful ones to have more of a say.

      > Users don't want their shit changing all the time.

      Yep that’s why you don’t ask them.

      I get if you have a specific flow that your use to. It would annoying for me too if that changed (as I’m pretty stubborn don’t like ppl making changes on my behalf), but that doesn’t mean it’s an objective better experience for all users or users who have yet to be familiar with the apps process.

      When these products operate in competitive markets and not some winner takes all market these are often about improving users experience.

      If this was something more high stakes like a medial trial I’d get it, but for stuff like filling out a document or watching a piece of media. The stakes for most SASS app are really low.

      1. jordwest · · focus · HN ↗
        > This is a bit hyperbolic, empirics is something that should be used more by decision makers

        In my experience working in software, empiricism is the only thing valued anymore. Intuition and thoughtfulness is out the window because it's not scientific enough. I would say most software now reflects that - it's almost all bland and statistically optimized to maximize engagement or revenue.

        > If this was something more high stakes like a medial trial I’d get it, but for stuff like filling out a document or watching a piece of media. The stakes for most SASS app are really low.

        There are plenty of subtle patterns used in SaaS form filling things too. For example notice that the "primary button" is always chosen as the one that will make the company the most money or collect the most data.

        Likewise, popups are annoying, but they result in more conversions. Forced logins are the same (how many form filling apps now force you to sign up with an account that you'll never use again, so that you can become a potential lead in future).

        Google recently started doing all of these on anonymous searches with a modal overlay and a big bright blue "Continue" primary button that takes you to a login screen, while a "don't sign in" button appears as far less noticeable text above it.

        It's at the point now where I'm surprised when any software gives you an option without blatantly telling you which one they want you to pick for their own benefit.

        A lot of it seems innocuous but I feel we're at the stage of death-by-a-thousand-cuts at this point.

        1. akst · · focus · HN ↗
          Hey Jordan hope you're doing well.

          > In my experience working in software, empiricism is the only thing valued anymore. Intuition and thoughtfulness is out the window because it's not scientific enough.

          With things like A/B testing, its not entirely an objective as you need to make assumptions which can be difficult to measure (although typically randomisation solves a lot of them), but you can only measure what you've decide to measure (which isn't random), so you don't know when you're in a local max. So IMO intuition and thoughtfulness is necessary. Sometimes product managers don't listen to data scientists when they say you can't measure Y with X, or the research design violates the required assumptions to make a causal claims (like reverse causality or controlling on a post treatment effect, e.g. employment as control when measuring income after hospitalisation (the treatment)). I think the worse offences I've seen have been from marketing teams.

          But proper research design does require intuition and thoughtfulness, because statistical models require thought, like other forms of supervised learning.

          I've seen both

          - PMs use questionable experiments to justify shipping something.

          - PMs dismiss experiments when it was a null result and shipped anyways.

          In either case I don't think the methodology is the cause of problems here, although I think shipping with a null result is justifiable if it's a larger unit of work (provided its not a regression).

          Sometimes things that have heterogenous effects get measured as a homogenous effect, Like say:

          - Your primary user base is X1 and X2 is a larger consumer base but makes up a small portion of your user base.

          - Your experiment does poorly with X1, but say there was an increase in user base X2.

          - However because X1 dominates the user base and your signups (because say you target ads to X1 over X2), no one drills into the effects on these different user bases, the result gets discarded as it seems to be a bad outcome.

          There's valuable information in the experiment outcome but without thought and attention you can miss it.

          > Google recently started doing all of these on anonymous searches with a modal overlay and a big bright blue "Continue" primary button that takes you to a login screen, while a "don't sign in" button appears as far less noticeable text above it.

          I mean that sucks, but IMO with their market share, the way Google chrome is inclined to develop their product is very different to firms in more competitive spaces.

          Perhaps A/B Tests, allows google to optimise the things they are incentivised to pursue, but in the hands of smaller firms with different incentives are willing to tweak things to be more appealing to users when they have far less market power, which I think is probably more the issue in the case of Google.

          I just don't think this is a universal problem with the methodology

          1. jordwest · · focus · HN ↗
            Oh hey didn't check your name before I replied! I assume this is the akst I know IRL, small world. Hope you are well too.

            > I just don't think this is a universal problem with the methodology

            In an isolated world I would agree with you, but in the messy reality we live in I think the broader problems with the methodology are two:

            1. The belief that everything can be measured. There are many intangibles (user trust, willingness to put up with bugs, "vibes") that aren't easily measurable. Yes, net promoter scores etc, but every company I can think of that uses them builds bland, buggy, largely disliked products that people use only because they have no choice.

            2. Focusing on experiments and measurable outcomes creates a tendency to de-prioritise anything that isn't easily measurable. For example larger, riskier projects that can't be quickly tested. Or whimsy - easter eggs that devs added to many products in the past that users remember for years. Rarely added anymore because they can't be justified against a roadmap full of experiments.

            Sometimes it feels a bit like we're sitting around so focused on measuring whether guests prefer one dish or the other and arguing over which experiment is best, that we don't notice a forest fire is blazing outside.

            To me maybe it's a bit of a question of science vs art. Take the video games industry for example, the AAA studios are largely moving toward the science end of the spectrum with predictable games that extract maximum engagement, while indie studios are mostly building small, unique experiences that don't try to dominate your attention.

            Experimentation always narrows focus down, and I think it's worth considering what's being traded off by doing so.

            1. akst · · focus · HN ↗
              It's possible I've spent too much time looking at town planning where nothing is really ever measured and when numbers are produced it is it some insane

              Like claiming townhouses not facing out into the street somehow produces X $ in mental health costs due to "lack of inclusion" and the footnote links to a study where non-english speaking communities in Australia were having real health costs due to lack of access to translation services. Which is something that passes for "evidence based" planning. Doing experiments in social sciences is a lot harder tho.

              Maybe it's too easy to do experiments in Software and like you said

              > Sometimes it feels a bit like we're sitting around so focused on measuring whether guests prefer one dish or the other and arguing over which experiment is best, that we don't notice a forest fire is blazing outside

              I do think they're a useful tool but I see what you're saying.

              Part of me feel some of that is risk aversion, but also maybe some of its process dependence when you have a number that provides strong certainty on a number of things, and because there's that feedback loop they get drawn to the things that cause the number go up. Similar to how people become dependent on LLMs to get stuff done or affirm if they did the right thing, and when they're in a space that's harder to measure they don't know to judge if they've done a good job or not.

              I'm spending less time working on software, as I decided to get an econ degree, specifically econometrics, so I do spent a lot of time trying to think how to better measure stuff specifically in public policy, so I might be bias lol

              But I do get what you're saying, hope things are well for you

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.