Tether’s AI research unit QVAC has released an open-source, cross‑platform BitNet LoRA fine‑tuning framework that allows 1‑bit large language models (LLMs) with up to and beyond 1 billion parameters to be trained and run on consumer hardware, including laptops, desktop GPUs and modern smartphones. The framework, integrated into QVAC Fabric and based on a modified llama.cpp stack, brings LoRA fine‑tuning and accelerated inference for Microsoft’s BitNet b1.58 architecture to heterogeneous GPUs such as Intel, AMD, Apple Silicon M‑series, and mobile GPUs like Adreno, Mali and Apple Bionic, without relying on Nvidia or cloud infrastructure.
Technically, the release demonstrates that BitNet’s 1‑bit (ternary-quantized) design, combined with LoRA adapters and QVAC’s Vulkan/Metal GPU backends, can sharply reduce memory and compute requirements compared with traditional 16‑bit or even 4‑bit quantized models. Benchmarks published by Tether show up to 77.8% less VRAM usage than a 16‑bit Gemma‑3‑1B model and the ability to fine‑tune models roughly 2× larger on edge devices than comparable Q4 non‑BitNet models, while achieving 2.1–11.3× faster inference on mobile GPUs versus CPUs on flagship phones such as Samsung S25, Google Pixel 9 and iPhone 16. In practical terms, users can fine‑tune a 125M‑parameter model in about 10 minutes and a 1B‑parameter model on a Samsung S25 or iPhone 16 in roughly 1–1.5 hours, with experimental runs up to 13B parameters on mobile devices.
Strategically, QVAC positions this framework as part of an “edge‑first” AI infrastructure that can decentralize AI training and inference, weakening dependence on large cloud providers and specialized data‑center GPUs. By enabling privacy‑preserving, local customization of LLMs on widely available devices, Tether argues that this approach could support large‑scale deployments of AI agents and applications such as federated learning, and broaden access to advanced AI capabilities in the Web3 and broader technology ecosystem.
✨ AI-generated background, compiled from web sources — not editorial content.