Keywords: NPU, unified RAM
Apple is doing it, AMD is doing it.
GPUs are an inefficient way of doing inference. They’ve been great for research purposes, into what type of NPU may be the best one, but that’s been answered already for LLMs. Current step is, achieving mass production.
5 years sounds realistic, unless WW3.
Powderhorn@beehaw.org 3 weeks ago
For the security tradeoff of sensitive data not heading to the cloud for processing? Not all businesses, but many would definitely see value in it. We’re also discussing this as though the options are binary … models could also be hosted on company servers that employees VPN into.