A new inference engine to run Kimi K3 2.78T parameter with 29GB of RAM marcobambini.substack.com 4 points by marcobambini 2 days ago
armchairhacker 2 days ago Current speed is “approximately one third of a token per second” marcobambini 2 days ago Right, we trade speed for the ability to run a 2.7T-parameter model while preserving accuracy. It is a first version, and we plan to improve the inference performance.
marcobambini 2 days ago Right, we trade speed for the ability to run a 2.7T-parameter model while preserving accuracy. It is a first version, and we plan to improve the inference performance.
Current speed is “approximately one third of a token per second”
Right, we trade speed for the ability to run a 2.7T-parameter model while preserving accuracy. It is a first version, and we plan to improve the inference performance.