‹ BackHN Continuity

Thread

How do we prevent mathemathics from devolving into the Medieval Era of secrecy?

158 points · 146 comments · jjgreen

  1. rramadass · · focus · HN ↗
    > Regardless of the true cost, it seems that professional mathematicians now need to wary about what they put into a LLM and think hard about how to disclose and publish a result.

    This is all but guaranteed now.

    Mathematicians/Scientists/Researchers need to stop sharing freely with "AI Companies" and have explicit clauses in place in their publications about not using their research without their explicit consent.

    There should be a clear legal distinction between using research data for AI model-training vs. another researcher using it.

    Come up with a legal framework, establish procedures for sharing and using others work and have a single scientific body in charge of enforcing it.

    1. nradov · · focus · HN ↗
      Just putting a clause in a publication won't prevent it from being used as training data. Information wants to be free.

      The frontier LLM vendors do sell enterprise licenses which contractually guarantee that your prompts won't be used for training. (Maybe they'll secretly violate the agreement but in principle it's legally enforceable.) Scholars and universities who care about credit and attribution will either have to purchase those licenses or run their own private open-weight LLM instances.

      1. oldsecondhand · · focus · HN ↗
        Even the $20 tier of ChatGPT has privacy settings that forbid using the user's data to be used for training. The question is, whether this setting is respected.
        1. tosapple · · focus · HN ↗
          the existence of the triplets NSA/CIA/GRU implies an imperitive no.
          1. r_lee · · focus · HN ↗
            that is a whole different thing though... AI labs are not government intelligence agencies
            1. tosapple · · focus · HN ↗
              minus two or three things, corporate data is classified as public/goverment data. we just saw something about earmarking domain last month? two being imminent domain. three being natsec.

              risk of prescient theory is more important than dismissive ablation.

              edit: to wit, facebook google and anything else not e2e.

              it's not like the ai is homomorphic.

              1. tosapple · · focus · HN ↗
                oh, haha. another point to make: what do you think they've been building out for 25 years with fusion centers and maryland/utah?

                "government intelligence agencies" ARE 'AI'.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.