I tried them, but could only test Pro none and Flash none and low, the other ones (medium/high) used way too many tokens and all requests timed out. Not sure if they have a problem with their API, or this model is really token inefficient/basically unusable.
The full-response APIs of many providers have been inadequate for a while now because their timeout intervals do not account for lots of thinking. You can use the streaming API to avoid timeouts.
XCSme · · focus · HN ↗
gpugreg · · focus · HN ↗
XCSme · · focus · HN ↗
16 minutes per test, which is a lot for simple questions...
Sometimes they fail because they reason more than their max context window without giving an answer, that's odd too.