Local models & providers

Built-in local inference (fully offline) or bring your own cloud provider — keys stay in the OS keychain.

Two kinds of models, switch anytime

ChengYoung supports two kinds of models. You can switch anytime in Settings, and use different models for different tasks:

Local modelsCloud providers
Data flowStays on this machine — inference runs offlineSent on demand to the provider you configure
CostFreeBilled by the provider
PrerequisiteDownload the model once (online)Your own API key

The local inference engine is built into the app (based on llama.cpp) with GPU acceleration for modern graphics cards:

  1. Choose "Local models" in Settings or the first-launch guide.
  1. Pick a model from the built-in curated list (mainly open-source series such as Qwen, across sizes — smaller models respond faster, larger ones are more capable).
  1. The model file downloads automatically on first selection: on Chinese networks the ModelScope mirror is preferred, otherwise Hugging Face; you can switch the download source in Settings.
  1. Once downloaded it works fully offline — chat, file access, and running commands all work without a network.

Model files are stored only on your computer; uninstalling the app does not delete them, and you can remove them from the model settings when cleaning up.

The local model picker

Connect a cloud provider

To use cloud models (stronger capabilities or longer context):

  1. In "Settings → Models" pick a cloud provider, or choose "OpenAI-compatible" to connect a custom endpoint.
  1. Enter that provider's API key.

Key safety: API keys are stored only in the operating system keychain (Windows Credential Manager) — never in plaintext config files, and never sent to any third party.

Keep in mind: with a cloud provider, your conversation content and related data are sent to that provider on demand and subject to its privacy policy — see Privacy & data.

Which to choose?

  • Privacy first / offline use: local models.
  • Complex long tasks: a more capable cloud model.
  • Daily mix: pin your frequent models in the switcher and change per task.

This page is adapted from upstream open-source documentation (Apache-2.0) with modifications; see /legal/provenance for sources and changes.