a lightweight hybrid reasoning MoE model with 7.9B total parameters and only 1.3B activated parameters per token. It is designed to deliver strong reasoning and agentic capabilities under a small inference compute footprint, making advanced model capabilities more accessible for local and resource-constrained deployment.

  • BeefAndPoultry@lemmus.org
    link
    fedilink
    English
    arrow-up
    0
    ·
    3 days ago

    my laptop is crappy, so like 5 tokens per second lol, prompt processing of like 20 tokens per second

    I think a decent laptop nowadays, even running CPU only, could probably do like 5x faster