Tiiny AI Pocket

Own Your AI in Your Pocket

Built to run your agents with zero friction, complete privacy, and zero ongoing fees.

W_0a55a2bc-94f5-4250-bf5f-6874d97cc50d.webp__PID:d6498ad4-837d-4dbd-bd09-2381f692d433
  • Run up to 120B in your pocket
  • Zero-config Open-source LLMs & Agents
  • 0 Token Fees
  • 30W Thermal Design power
  • Bank-grade security for sensitive work
  • 1TB NVMe PCIe 4.0 SSD
  • SoC (Armv9.2 CPU + NPU) 30TOPS + dNPU 160TOPS

Boost Your PC in 3 Steps

  • Plug in

  • Pick your model and agent

  • Start working

The Smallest Edge-AI Device for Local LLMs

Enables server-level AI without extra GPUs or PC upgrades, running 120B models on a phone-sized AI device to boost productivity anywhere.

35.webp__PID:5dbdfd09-2381-4692-9433-9b32c4fde48a

See the Real Performance

Tested with 512-token inputs generating up to 2048 tokens output (max token limit 2048).

09_ca4c823e-0478-4b6d-b594-16da923af321.webp__PID:f692d433-9b32-44fd-a48a-e9da4d594d09

PowerInfer: Faster Inference on Consumer GPUs

PowerInfer v1 was the first infra to run 175B-parameter models on a consumer RTX 4090, reaching 90% of A100 performance and up to 11.69x faster inference than prior methods. the new generation we adopt pushes even further.

10-PowerInfer.webp__PID:702966cd-64cb-4463-8763-289b9c3bf4e4

Open-Source Ecosystem,Unlocked

One-click deployment, no technical setup required.Continuous SOTA open-source LLMs and hardware-levelOTA updates.Already building with APIs? Use Tiiny as a local token fact
ory with OpenAI API compatibility. Infinite tokens, zero token fees.

横版Agents(640P.gif__PID:3f9be7d0-62da-4a79-8f2b-c1b9e35cd9f8

Bank-Grade Security

All data stays unreadable if the drive is removed. Safe for proprietary code, research, and personal knowledge.

  • Hardware-level AES-256 full-disk encryption

  • User-controlled keysin a secure enclave

  • 100% on-device processing
14-隐私-详情.webp__PID:de0ccf67-b8c9-446f-8933-822c7136207a

True Long-Term Memory, Truly Knows You

16_8fa6ba54-e238-49b3-907d-5a36739136a6.webp__PID:c4fde48a-e9da-4d59-8d09-2f5a39481dae

Comprehensive Developer SDK

Build, extend, and customize local Al applications and Al-native devices.

17-SDK.webp__PID:7e13d61b-723e-4afd-a7ac-388c75e39486

Own Your Al, Once and for All

Tiiny AI Pocket pays for itself fast- while keeping data local.

预热页0402-1.webp__PID:016708ec-43ac-4437-8389-bff39f88a0c3

Specification

LLMs: GPT-OSS-120B, GPT-OSS-20B,Llama3.1-8B-Instruct, Gemma3-4B,Ministral-3-8B-Instruct-2512,Qwen3-30B-A3B-Instruct-2507,Z-Image-Turbo,Qwen3-Reranker-0.6B, etc.

Avg. Output Speed: 18-40tokens/s

SoC: CPU (arm v9.2) + NPU, 30 INT8 TOPS

dNPU: 160INT8 TOPS


Memory: 80GB LPDDR5X @6400MT/s

Storage: 1TB PCIe 4.0 SSD


Audio: Internal mono audio output

Microphone: Built-in digital microphone

Bluetooth: BT5.3w/LE

Wi-Fi: 802.11a/b/g/n/ac/ax, Wi-Fi max transfer speed: 2.4 Gbps

Interfaces: Type-CX3

Power: TDP 30W (65W power adapter required)


Weight: 300g

Dimension :142x80x22 mm


Compatible System: macOS & Windows

When you sign in with your Google account, we use your basic profile (name, email address and profile picture) solely to create and access your Tiiny account. We do not access your Gmail, Google Drive, Google Calendar or any other Google services, and we never share your Google user data with third parties.