Skip to content

Ask Ying-Ja Chen if they know about Multi-Process Service to lower the VRAM overhead, potentially having it run on any modern GPU architecture #27

Description

@keiran-rowell-unsw

Early indications is that typical query sizes will not fully saturate everything and it's mostly about start-up cost, or idle time if running in GPU-server mode.

Would need to bundle together user request batches to really get the scale.

Mentor shared that the NIMs already use MMSeqs-GPU in the background (including AF2!), but not prior experience with Multi-Process Service so that idea was likely idealistic

Metadata

Metadata

Assignees

No one assigned

    Labels

    enhancementNew feature or request

    Type

    No type

    Fields

    No fields configured for issues without a type.

    Projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions