Arm
2 updates on Arm.
FlexServe runs on-device LLM inference inside TrustZone for 4.41 percent more latency
Shanghai Jiao Tong University researchers kept model weights and prompts inside ARM TrustZone at 4.41 percent more time to first token.
Arm runs Llama 2 7B at 9.6 tokens per second on three mobile CPU cores
Arm demonstrated Llama 2 7B generating 9.6 tokens per second on three Cortex-A700 CPU cores, using int4 kernels it says beat llama.cpp by 20 percent.