Kynora is a private, offline-first AI workspace for macOS, iPadOS, and iPhone, currently in development. It will let you chat with a local AI model and build a private, searchable knowledge base from your own documents — all running entirely on your device, with no cloud servers involved.
Ask Docs will be retrieval-augmented Q&A over documents you import — answers grounded in your knowledge base with a Sources panel (title, heading, page label, excerpt). Chat will be general multi-turn conversation with a local model, without document retrieval. Each mode is designed to use its own default generation model.
Kynora is still being built, and we don't have a fixed release date yet. Check back here for updates as development progresses.
Yes. Kynora is designed to run all AI inference, embedding, and search locally using Apple's MLX framework on Apple Silicon, or llama.cpp elsewhere. Once the app and a model are installed, it's built to work fully offline — even in airplane mode.
Cloud AI tools send documents to external servers for processing. Kynora is being built to run parsing, embedding, retrieval, and generation on your device instead — your files, queries, and conversation history are designed to never be sent for cloud inference, and no cloud AI account will be required for private Chat or Ask Docs.
The plan is PDF, DOCX, TXT, and Markdown at launch, with scanned pages handled through on-device OCR. Additional formats are being considered for later releases.
You'll be able to choose open-source models from curated GGUF and MLX catalogs for generation, embedding, reranking, and vision. Downloads will go from the model provider directly to your device. You'll also be able to import local weights or optionally use a local Ollama server for Chat. Nothing large is planned to be bundled by default.
Kynora is being built for macOS, iPadOS, and iPhone. Apple Silicon is recommended for the best on-device AI performance, but Intel Macs are also supported through llama.cpp.
There's no Kynora cloud inference backend planned. Document parsing, chunking, embedding, hybrid search, optional reranking, and LLM inference are designed to run entirely on your device. Model downloads will go from Hugging Face (or similar) directly to your device — not through a Kynora server. The app is being built without analytics or telemetry.
Pricing hasn't been finalized. What we can say now: there's no cloud AI subscription required to use it, since inference runs on your own device. Any pricing details will be shared here before launch.