While this is great, the training is where the compute is spent. The news is also about R1 being able to be trained, still on an Nvidia cluster but for 6M USD instead of 500
you are viewing a single comment's thread
view the rest of the comments
view the rest of the comments
replies: