Early indications is that typical query sizes will not fully saturate everything and it's mostly about start-up cost, or idle time if running in GPU-server mode.
Would need to bundle together user request batches to really get the scale.
Mentor shared that the NIMs already use MMSeqs-GPU in the background (including AF2!), but not prior experience with Multi-Process Service so that idea was likely idealistic
Early indications is that typical query sizes will not fully saturate everything and it's mostly about start-up cost, or idle time if running in GPU-server mode.
Would need to bundle together user request batches to really get the scale.
Mentor shared that the NIMs already use MMSeqs-GPU in the background (including AF2!), but not prior experience with Multi-Process Service so that idea was likely idealistic