Everything You Need

Tiiny AI Pocket running beside a laptop

Runs 24/7 with 0 token fees

Leave AI running in the background, without counting tokens or watching the bill.

Tiiny AI Pocket used offline while working

Pocket-sized and fully offline

Your AI goes wherever you go, from long flights to remote field work.

Tiiny AI Pocket connected to a laptop

Agents ready out of the box

Start with the task, not the setup. Pick an agent and get to work in seconds.

Tiiny AI Pocket protected inside a secure vault

Have Powerful AI, Keep Your Privacy

Everything runs on-device and under your control. Handle sensitive data confidently.

Instant AI Upgrade in 3 Steps

Ready to run in seconds, no complex setup.

  1. Plug in
  2. Pick your model and agent
  3. Start Working
Tiiny AI Pocket connected and ready to use

Up to 120 Billion Parameters
Fits in Your Palm
Up to 120 Billion Parameters
Fits in Your Palm

Models once reserved for servers are now brought to your pocket, driven by PowerInfer. Powerinfer v1 was the first infra to run 175B-parameter models on a consumer RTX 4090, reaching 90% of A100 performance and up to 11.69 times faster inference than prior methods. The new generation we adopt pushes even further.

Models once reserved for servers are now brought to your pocket, driven by PowerInfer. Powerinfer v1 was the first infra to run 175B-parameter models on a consumer RTX 4090, reaching 90% of A100 performance and up to 11.69 times faster inference than prior methods. The new generation we adopt pushes even further.

Tested, Not Claimed
  • 4K40Decode tok/s
  • 64K30Decode tok/s
  • 128K20Decode tok/s
  • 3000Prefill tok/s
  • 2000Prefill tok/s
  • 1000Prefill tok/s
TiinyOS

Made Intuitively Simple

Agent Store

Straight to Work with Zero Config Straight to Work with Zero Config

Skip the setup. Pick an agent, handling everything from vibe coding and data analysis to automation.

Skip the setup. Pick an agent, handling everything from vibe coding and data analysis to automation.

    Tab image
    Tab image
    Tab image
    Tab image
    Tab image

    As Easy as Your Favorite Apps

    Tiiny chat interface

    The Experience You
    Already Know

    Ask questions, upload files, and get instant answers—with the smoothness of a cloud AI service.

    Tiiny web search and connectors

    Bridge local compute
    with live world

    Optionally connect to the web for real-time information retrieval and dynamic fact-checking.

    Tiiny task execution and progress tracking

    Complex problems solved
    in one conversation

    Simply state your goal. Tiiny coordinates the right agents to handle the execution with step-by-step status tracking.

    Connectors

    Get 24/7 autonomous assistance

    Link your daily tools to automate replies, manage schedules, and get briefings. Leverage your existing workflows via native MCP support anytime.

    Scheduled Automation

    Set it and forget it

    Easily schedule recurring tasks to run in the background, freeing you for what matters.

    Tiiny Vault

    Own Your AI Brain

    The private knowledge base built exclusively for you, with all data stored on-device. Preset your preferences and let it evolve with every interaction via manual edits, live chat requests, or automated summaries.

    Tiiny Vault Interface

    No Rebuilding, No Lock-in

    Hugging Face and custom model import

    Bring Your Own Models

    Directly import from Hugging Face to experience global AI breakthroughs the moment they release. Or convert your custom models seamlessly.

    Unified compatibility across AI tools

    One Switch, Full Compatibility

    Drop-in replacement for OpenAI, Anthropic, and Ollama APIs. Seamlessly swap the backend of your existing tools without rebuilding anything.

    TiinySDK architecture with application framework, Tiiny Core and hardware abstraction layer

    TiinySDK: Build Your Own Intelligence

    Customize and extend Tiiny AI Pocket to your exact needs—as well as creating without limits, from bespoke AI agents to next-gen AIoT products.

    Your Data
    For Your Eyes Only

    Device access is solely yours. All data becomes instantly unreadable if the drive is removed. Zero-compromise security for proprietary code, research, and knowledge.

    Tiiny AI Pocket protected by an on-device security shield
    • 100% on-device processing.
    • AES-256 full-disk encryption.
    • User-controlled keys.

    Stay Chill, Outrun Boundaries

    Up to 3x Energy Efficiency

    Let every single watt deliver 3x tokens per second. Tiiny's 30W TDP slays the power-hungry monster of local LLM deployments. Same intelligence, stress-free running.

    Tiiny AI Pocket

    Power
    30W
    Tokens per second per watt
    0.5–1.4
    Annual CO₂ Emissions
    ≈21kg

    Others

    Power
    100–150W
    Tokens per second per watt
    0.29–0.58
    Annual CO₂ Emissions
    ≈45kg

    * Note: All performance data is based on comparative testing under a 4K context window length.

    Heavy Workloads, Silently Handled Heavy Workloads, Silently Handled

    Tiiny AI Pocket delivers smooth performance at under 35 dB, made possible by an ultra-thin vapor chamber, dual fans, and an integrated fin-and-fan cooling design engineered to eliminate localized heat.

    Tiiny AI Pocket delivers smooth performance at under 35 dB, made possible by an ultra-thin vapor chamber, dual fans, and an integrated fin-and-fan cooling design engineered to eliminate localized heat.

    Packed Tight,
    Travel Light

    Exploded view of Tiiny AI Pocket showing its enclosure, dual fans, cooling assembly and internal circuit boards

    Specifications

    SoC
    CPU (arm v9.2) + NPU 30 INT8 TOPS
    dNPU
    160 INT8 TOPS
    Memory
    80GB LPDDR5X @6400MT/s (32GB on SoC, 48GB on dNPU)
    Storage
    1TB PCIe 4.0 SSD
    Audio
    Internal mono audio output
    Microphone
    Built-in digital microphone
    Bluetooth
    BT 5.3 w/LE
    Wi-Fi
    802.11a/b/g/n/ac/ax (Wi-Fi max transfer speed: 2.4 Gbps)
    Interfaces
    Type-C × 3
    Power
    TDP 30W (65W adapter required)
    Weight
    300g
    Dimensions
    142 × 80 × 22 mm
    Compatible System
    macOS and Windows

    FAQ

    What does "120B MoE model" mean?
    MoE (Mixture of Experts) is the dominant architecture for large language models in 2024-2026. GPT-4, Mixtral 8×22B, DeepSeek-V3, and gpt-oss-120b all use MoE.
    In a 120B MoE model, the total parameter count is 120 billion — this determines the model's knowledge capacity and reasoning depth. However, during each inference step, only 5.1 billion parameters are activated (the "experts" most relevant to the current token). This is why a 300g device can run it efficiently.
    This is the same principle that allows GPT-4 to be powerful yet serveable — large total knowledge, efficient per-token compute.
    How fast is inference on Tiiny?
    Inference speed depends on model size and context length. Our measured benchmarks:
    • 20B models: ~25-40 tokens/second
    • 70B models: ~14-23 tokens/second
    • 120B MoE models: ~10-20 tokens/second
    These speeds are sufficient for real-time conversation and most agent tasks. TurboSparse acceleration further improves throughput on sparse models.
    Why is the memory split into two pools?
    The 80GB memory is divided into a 32GB main pool (high-bandwidth, for active model inference) and a 48GB extended pool (for model storage and context caching). This dual-pool design was chosen for three reasons:
    1. Thermal management: A single 80GB high-bandwidth pool would exceed the thermal envelope of a 300g fanless device.
    2. Cost efficiency: The extended pool uses lower-bandwidth memory, reducing cost without impacting inference speed (model weights are pre-loaded into the main pool before inference).
    3. Portability: ARM-based architecture with dual-pool design keeps power consumption at 15-25W, enabling USB-C PD operation.
    Is Tiiny based on open-source software?
    Yes. Tiiny's inference engine is based on PowerInfer, an open-source project developed at Shanghai Jiao Tong University (SJTU) with 9,100+ GitHub stars. Tiiny's engineering team has contributed significant optimizations including:
    • ARM architecture adaptation and optimization
    • TurboSparse sparse inference acceleration engine
    • MoE model inference optimization
    • On-device deployment engineering
    We believe in standing on the shoulders of giants — just as Android built on Linux, Tiiny builds on PowerInfer. We are committed to contributing our optimizations back to the open-source community.
    What APIs does Tiiny support?
    Tiiny provides full API compatibility with:
    • OpenAI API — drop-in replacement for GPT models. Point your existing code to Tiiny's local endpoint and it works.
    • Anthropic API — compatible with Claude API interface.
    • Ollama API — full compatibility with the Ollama ecosystem.
    • MCP Protocol — Model Context Protocol support for advanced agent workflows.
    This means you can replace cloud API calls with local Tiiny calls by changing a single URL — zero code rewrites needed.
    What is your privacy and data security model?
    Tiiny is designed for zero-data-leakage AI processing:
    • 100% offline: No network connection required. No telemetry, no analytics, no phone-home.
    • AES-256 hardware encryption: Full-disk encryption at the hardware level.
    • Data isolation: All inference, context, and conversation history stays on-device.
    • No accounts required: Use Tiiny without creating any account or sharing any personal information.
    This makes Tiiny suitable for GDPR, HIPAA, and attorney-client privilege scenarios where cloud AI cannot be used.

    Cart

    loading