Using the app

MoE maps

The technology the app is named after: running Mixture-of-Experts models larger than your computer's memory, by keeping ready the "experts" you actually use.

What's there

The app uses maps, has a page of their own (see The MoE maps page) and, with Expert mode, creates them on your own texts (see Creating a map). Coming: the portal to share them.

The idea

A Mixture-of-Experts model has tens or hundreds of experts in each layer, but for each word it uses only a few. And it doesn't use them all equally: on a certain kind of text, a small group of experts works almost all the time, the others rarely.

  1. Profile. The app has the model read texts like those you'll use it on (your documents, your code) and notes which experts activate.
  2. Map. The result is a map: which experts to always keep in memory and which to leave on disk.
  3. Load smartly. The "hot" experts stay in memory; the others are read from disk only when needed. This way a model larger than the PC's memory becomes usable.

Using a map

A map is a .moeplan file made for one specific model file. Put it in the plans folder inside the app's data folder, or next to the model. When you load the model, the app finds the map by itself; with Use the expert map on (the default, in Settings) it:

  1. locks the hot experts in RAM: the status shows Locking experts in RAM…. Tens of GB are read from disk: from a few seconds to a couple of minutes;
  2. loads the model, which finds those experts already in memory. When it is ready the status shows Ready · map active.

A map only applies to the exact file it was made for (same name, same size): another quantization of the same model needs its own map. The app never locks more than your total RAM minus the RAM to leave free set in Settings (12 GB by default, from 512 MB up). If there are several maps it picks the largest that fits, so keeping maps of different sizes lets it step down when you ask for more free RAM. If locking fails, the model starts anyway without the map, only slower, and Settings tells you why.

Selective fidelity

For each word the model chooses a few experts; sometimes one of them is outside the map but matters little to the answer. Selective fidelity uses the best expert in the map in its place, so the disk is read less often.

SettingWhat it does
OffThe model always uses the experts it chooses. Slowest.
Balanced (default)Replaces an expert outside the map only when it is at the edge of the model's choice and an expert in the map is almost as good. More than twice as fast on text similar to the map's; quality almost unchanged.
FastExperts in the map count four times as much when the model chooses. Up to about 2.6 times as fast on text similar to the map's, but on different text the quality drops noticeably. Best with a map made on your own content.

In our tests with Qwen3-235B on a PC with 64 GB of RAM and a 40 GB map, on text similar to the map's, generation went from about 1.4 words per second with the map alone to 3.1 with Balanced and 3.6 with Fast. On different text both reach about 2.

The MoE maps page

In the bar on the left, MoE maps shows every installed model with its maps: how much RAM they lock, how much they cover and their state (In use now, Used on the next load, Ready, or Doesn't fit in memory with the free RAM you chose).

  • Map to use: with several maps for the same model, choose which one to use; Automatic (the default) takes the largest that fits in memory. A choice that doesn't fit is never forced. If the model is already loaded, the new map applies from the next load: the page says so, with Reload now.
  • Import a map… checks that the file is a valid map and copies it into the plans folder; Open the maps folder opens it in File Explorer. Maps in that folder can also be deleted, after a confirmation.
  • Maps in the folder made for models you haven't installed appear separately, under Other maps in the folder.

When you select a map, on the right you see:

  • Expected and actual: the coverage the map expects next to the actual one, that is how many experts the model really finds in memory while you work. It updates live when the model runs with that map. If the actual value stays below the expected one, you are using different experts from the map's: a map made on your own texts would do better.
  • Coverage by memory: how much of the work the most used experts cover depending on the memory you give them, with the map's dot. It helps choose the size: past a certain point, more RAM adds little.
  • Expert heat: one row per layer. In each row, first the experts kept in memory, from the most used (violet) to the least used (cyan), then the ones left on disk, dark: the coloured part is as long as the number of that layer's experts in the map, so you see at a glance where the map concentrates.

The speed with and without a map, day by day, is in Statistics.

Creating a map

With Expert mode on (Settings), the MoE maps page shows Create a map…. A four-step procedure:

  1. Model. Only installed Mixture-of-Experts models appear: dense models don't need a map.
  2. Texts. Add folders or files similar to what you'll use the model on: PDF and Word documents, texts, code. The app takes about 80 KB, sampled across everything, and leaves out generated files, dependencies, binaries and files that are too large. Keys, tokens and passwords found in the texts are replaced before profiling: the app tells you which files they were in, without showing their value.
  3. Profiling. The model reads the text and the app notes which experts it picks for every word. A bar shows the tokens read and the time left; Stop halts everything, and the partial profile can already be used, with less text.
  4. Map. Choose with a slider how much RAM to give it, looking at the coverage/memory curve and the expected coverage. The proposed value is the largest that fits while leaving free the RAM chosen in Settings. Give it a name and press Create the map: it goes into the maps folder and, if you leave it selected, the model uses it from the next load.
While profiling, the PC is busy

On a model larger than RAM, profiling takes from a few minutes to half an hour (on our PC, about 2.5 minutes for 2,000 tokens with Qwen3-235B). The CPU runs flat out, the disk reads the whole model, free RAM fills with the model's cache and part of the VRAM holds the attention. The loaded model is unloaded before starting and the chat can't be used until it finishes.

The collected text stays on your PC and is deleted when profiling ends, whatever happens. The .moep profile stays next to the map with the same name: it lets you remake the map in another size without profiling again, and it doesn't contain the texts. A map in use keeps its RAM locked for as long as the model is loaded.

Map files

FileWhat it is
.moeplanThe map: which experts to keep in memory, in what order and where they are in the model file. It is a text file made for one specific model file (same name, same size), and the only one the app needs.
.moepThe profile: which experts the model chose while reading the sample texts. It is only used to generate maps and can be reused to make maps of different sizes.

Both are HotMoE formats. Neither contains the texts used to create them, only statistics about the experts: sharing a map doesn't reveal your documents.

Coming soon

Share and download maps from a community portal, only if you want to.

HotMoE is a Virsion project. This guide describes the app as it is today: what is still on the way is marked as such.