• Septimaeus@infosec.pub
    link
    fedilink
    English
    arrow-up
    0
    ·
    17 hours ago

    Atm the meta for local inference is unified memory and “routed” local agents (multiple smaller role-specific agents in a trench coat)

    The former is standout for cost efficiency (e.g., 4x RDMA 48gb Minis for a 192gb cluster @ $43/gb vs a $45k b200 alone)

    The latter is standout for many things, including resource efficiency on smaller machines