FreeToken: treating a personal machine as an elastic platform for serving large MoE models
FreeToken rethinks local inference around two realities of edge AI: agent workloads keep changing shape, and every machine's resource balance differs. It co-designs model layout, expert residency, CPU-GPU execution and memory management instead of committing to a fixed offloading strategy.
- 348 stars/7d
- 2 citations
- 109 upvotes