We don’t know the model architecture but there’s a lot of evidence to suggest load dependency on the hardware affect models quality (see for example how the original Google Translate models got worse depending on time of day). The GPUs at maximal utilization is what they’re shooting for with their pricing models so you’d need to measure model performance at peak loads to know the floor of performance I would think.
kingcauchy · · focus · HN ↗