LocalAI: Offline AI Chat LLM

LocalAI: Offline AI Chat LLM

Screenshots

Screenshot 1
Screenshot 2
Screenshot 3
Screenshot 4
Screenshot 5
Screenshot 6
Screenshot 7
Screenshot 8
Screenshot 9
Screenshot 10
Screenshot 11
Screenshot 12
Screenshot 13
Screenshot 14
Screenshot 15
Screenshot 16
Screenshot 17
Screenshot 18
Screenshot 19
Screenshot 20
Screenshot 21
Screenshot 22
Screenshot 23
Screenshot 24
Screenshot 25
Screenshot 26
Screenshot 27
Screenshot 28
Screenshot 29
Screenshot 30
Screenshot 31
Screenshot 32

Description

LocalAI: Your 100% Offline, Private AI Assistant

Are you concerned about sharing private data, confidential documents, or personal chats with cloud-based AI? Do you need a powerful AI assistant when traveling, commuting, or working in areas without internet access?

Transform your Android device into a secure, private AI workstation with LocalAI. LocalAI is a 100% offline AI chatbot that runs open-weight Large Language Models (LLMs) entirely on-device. There is no cloud processing, no mandatory subscriptions, and absolutely zero data collection. Your prompts, documents, chats, and photos never leave your phone.

🚀 Why Choose LocalAI?
• 100% Private & Secure: LocalAI processes everything locally on your hardware using our highly-optimized, crash-resistant Llama.cpp inference engine. No telemetry, no backend tracking, and no internet required.
• Complete Data Ownership: Total peace of mind. Analyze confidential work, legal PDFs, and personal journals securely offline.

🧠 Run State-of-the-Art Open-Source LLMs
Discover, download, and manage GGUF models directly within our Hugging Face-powered Model Hub. Supported cutting-edge architectures include:
• Meta LLaMA 4 (Scout, Maverick) & Llama 3.x
• Google Gemma 4 & Gemma 4 Mobile
• DeepSeek-V4 (Flash & Pro distilled)
• Alibaba Qwen 3.5 & Qwen 3.6 (including MTP architectures)
• Ornith 1.0 & Bamboo 1 (high-efficiency models)
• IBM Granite 4.1 & Microsoft Phi-4

⚡ Hardware-Accelerated Local Inference
Designed for extreme speed and memory efficiency:
• Flash Attention 2: Hardware-accelerated attention for faster token generation.
• KV Cache Quantization: Q4_0/Q8_0 caching saves 30-40% RAM, preventing Out-Of-Memory crashes.
• GBNF Grammar & JSON Schema: Force structured outputs natively.
• Response Telemetry: Real-time tokens/sec, prompt/sec, and RAM/CPU hardware monitors.
• Reasoning Support: Natively surfaces `` reasoning blocks (DeepSeek-V4, Ornith).

📄 Chat with PDFs and Documents (Offline RAG)
Import PDFs, Word (.docx), Excel (.xlsx), CSVs, or text files. LocalAI parses, chunks, and embeds content locally using on-device Vector RAG (sqlite-vec). Summarize, ask questions, and chat with PDFs offline securely without an internet connection.

🖼️ Multimodal Vision AI Offline
Load any vision-capable model (like SmolVLM, LLaVA, or Qwen-VL) to chat about your photos. Take a picture or import an image to summarize, extract text, or analyze layouts—processed 100% offline.

✨ Artifact Mode (Interactive UI Generation)
Auto-promotes LLM-generated code blocks (HTML, SVG, Python, etc.) into interactive Sandboxed UI cards. Render UI mockups, charts, or games directly inside your private chat!

☁️ BYOK (Bring Your Own Key) & Hybrid Cloud
Need more power? Upgrade to Premium to switch between on-device and cloud models:
• BYOK API Integrations: Connect to OpenAI (GPT-4o), Anthropic (Claude 3.5 Sonnet), or OpenRouter using your own API keys.
• Custom Endpoints: Connect to self-hosted Ollama, vLLM, or local servers on your home network.
• Advanced Web Search: Scrape up to 10 live web results for real-time answers.
• Ad-Free Experience.

🎨 Ultimate Customization & UI
• 19 Themes: Nautilus, Cyber, Aura, Monokai, Sunset, Emerald & more.
• 12 Backgrounds: Circuit, Matrix, Dots, Grid, Topography & more.
• 18 Fonts: Inter, Roboto, Poppins, FiraCode, EBGaramond & more.
• 38 Languages: English, Spanish, French, German, Chinese, Hindi, Japanese & more.

💬 SQLite Local Chat History
Your chats are saved locally on your device with full Markdown rendering, LaTeX math formulas, zero-dependency syntax-highlighted code, and quick-copy buttons.

Note: Local performance is hardware-dependent. Devices with 8GB+ RAM and modern high-end processors (e.g., Snapdragon 8 Gen 3+, Dimensity 9300+) will experience significantly faster token-per-second generation.

Developer

Apex Creators

Developer ID: 8845231963384448407

View More Apps →

Recent Reviews

Jack Wang

v26.6.5

还不错,能不能编写一套OpenCL支持GPU加速?

Jul 7, 2026

Get API Access to 1.5 Million Apps

Access complete app metadata, rankings, reviews, and screenshots for your product and market intelligence workflows.