So predicting the next word given all humanity’s knowledge is surely going to max out at slightly less good (we probably can’t get perfect data) than the best human in any specific field. What test does the AI do to be able to understand it is improving? At some point it becomes impossible to know that the output is actually better right?
Yup, the e2e proof of FLT (estimated effort of 5 years & 1M$ by the best human in the field) and a counter example for a millennium prize (similarly valued at 1m$.
andy_ppp · · focus · HN ↗
tucnak · · focus · HN ↗
Citation needed
andy_ppp · · focus · HN ↗
NitpickLawyer · · focus · HN ↗