Skip to content

Choose a smart search model

Smart search uses a CLIP model to understand what’s in your photos. The default, ViT-B-16-SigLIP-384__webli, is fast and light, and works well for English searches. Other models can give better results or understand other languages, at the cost of more memory and slower processing, both when your library is indexed and every time someone searches.

This page is for administrators. Changing the model means reprocessing every photo and video, so choose once, carefully.

The first question is which languages people will search in.

  • Only English: start with the English models below.
  • Mostly one other language: an nllb model usually gives the best results.
  • A mix of languages: an xlm or siglip2 model is more flexible.

There are two kinds of multilingual model:

Model family How it reads a search
nllb Expects the search to be in the language set in the user’s own settings
xlm and siglip2 Understands the search whatever language it’s in, regardless of the user’s language setting

nllb models tend to perform best, and are the better choice when people mainly search in their own, non-English language. xlm and siglip2 models suit people who switch between languages from one search to the next.

  1. Copy the model name, for example ViT-B-16-SigLIP__webli.
  2. Go to Administration, then Settings, then Machine Learning Settings, then Smart Search.
  3. Paste the name into CLIP model.
  4. Save the settings.
  5. Go to Administration, then Jobs.
  6. Click All next to Smart Search to reprocess your library with the new model.
  7. Optionally, check the logs of the server and the machine learning container for errors.

The model downloads the first time it’s used, which can take a few minutes for the larger ones. Until the Smart Search job finishes, smart search results are incomplete.

In rare cases, switching can leave some of the old model’s data behind and cause errors in Smart Search jobs. If you see that in the logs, switch back to the previous model and save, then repeat steps 3 to 7.

Changing the model also affects features built on smart search. The labels behind suggested tags and the CLIP queries in smart albums are encoded again with the new model automatically. The search data for video moments is cleared, while the frames and moments themselves stay.

The figures below come from benchmarks run without acceleration on a single desktop processor, at full (f32) precision, which is Frameleaf’s default. Treat them as a way to compare models, not as what your server will use.

Column Means
Memory (MiB) Peak memory used by the model, not counting image decoding, parallel jobs or the web server
Time (ms) Average time for one pass of the model once it’s warmed up
Recall (%) How often the right photo appears in the top results, averaged over standard test sets. Higher is better.

A model is optimal for a language when no other model beats it in one respect (memory, time or recall) without being worse in another. Prefer optimal models for the languages that matter to you.

These are the optimal models for English searches, best results first. Larger models near the top understand detailed and unusual searches better; smaller ones further down are much faster, and not that different in quality.

Model Memory (MiB) Time (ms) Recall (%)
ViT-SO400M-16-SigLIP2-384__webli 3854 56.57 85.99
ViT-L-16-SigLIP2-512__webli 3358 92.59 85.75
ViT-SO400M-16-SigLIP2-256__webli 3611 27.84 85.62
ViT-SO400M-14-SigLIP2__webli 3622 27.63 85.53
ViT-L-16-SigLIP2-384__webli 3057 51.7 85.47
ViT-L-16-SigLIP2-256__webli 2830 23.77 85.03
ViT-B-16-SigLIP2__webli 3038 5.81 84.86
ViT-B-16-SigLIP-512__webli 1828 26.17 83.28
ViT-B-16-SigLIP-384__webli (default) 1128 13.53 83.19
ViT-B-32-SigLIP2-256__webli 3061 3.31 82.28
ViT-B-16-SigLIP__webli 1081 5.77 81.9
ViT-B-32__laion2b-s34b-b79k 1001 2.29 77.62
ViT-B-16__laion400m_e32 975 4.98 76.43
ViT-B-32__laion400m_e31 999 2.28 73.83
ViT-B-32__openai 1004 2.26 69.9
RN50__openai 913 2.39 69.02
RN50__cc12m 914 2.37 64.59
RN50__yfcc15m 908 2.34 53.63

For each language tested, this table shows the model with the best recall, and the best optimal model that needs about 3 GB of memory or less. The figure in brackets is the recall for that language.

Recall varies a lot between languages, and the default model is much weaker outside English and a few European languages. If people search in a language below, a different model will usually make a big difference.

Language Best results Best at about 3 GB or less
Arabic nllb-clip-large-siglip__mrl (77.3) ViT-L-16-SigLIP2-384__webli (68.25)
Bengali nllb-clip-large-siglip__v1 (76.16) ViT-B-16-SigLIP-i18n-256__webli (36.43)
Chinese (Simplified) nllb-clip-large-siglip__v1 (79.7) ViT-L-16-SigLIP2-384__webli (71.11)
Croatian nllb-clip-large-siglip__mrl (87.46) XLM-Roberta-Base-ViT-B-32__laion5b_s13b_b90k (72.69)
Cusco Quechua nllb-clip-large-siglip__mrl (38.08) None listed
Czech nllb-clip-large-siglip__mrl (73.76) ViT-L-16-SigLIP2-384__webli (62.12)
Danish nllb-clip-large-siglip__v1 (87.16) ViT-L-16-SigLIP2-384__webli (79.82)
Dutch ViT-SO400M-16-SigLIP2-512__webli (80.05) ViT-L-16-SigLIP2-384__webli (79.49)
Filipino nllb-clip-large-siglip__mrl (67.57) ViT-B-16-SigLIP-i18n-256__webli (36.81)
Finnish nllb-clip-large-siglip__mrl (84.27) ViT-B-16-SigLIP-i18n-256__webli (63.16)
French ViT-SO400M-16-SigLIP2-384__webli (86.5) ViT-L-16-SigLIP2-384__webli (85.35)
German ViT-SO400M-14-SigLIP2-378__webli (87.32) ViT-L-16-SigLIP2-384__webli (86.56)
Greek nllb-clip-large-siglip__mrl (74.58) XLM-Roberta-Base-ViT-B-32__laion5b_s13b_b90k (64.69)
Hebrew nllb-clip-large-siglip__v1 (88.04) ViT-B-16-SigLIP-i18n-256__webli (74.59)
Hindi nllb-clip-large-siglip__mrl (62.02) ViT-L-16-SigLIP2-384__webli (35.76)
Hungarian nllb-clip-large-siglip__mrl (85.59) ViT-B-16-SigLIP-i18n-256__webli (73.95)
Indonesian nllb-clip-large-siglip__v1 (85.46) ViT-L-16-SigLIP2-384__webli (84.58)
Italian ViT-SO400M-16-SigLIP2-512__webli (87.17) ViT-L-16-SigLIP2-384__webli (86.35)
Japanese XLM-Roberta-Large-ViT-H-14__frozen_laion5b_s13b_b90k (83.95) XLM-Roberta-Base-ViT-B-32__laion5b_s13b_b90k (75.93)
Korean nllb-clip-large-siglip__mrl (80.56) ViT-L-16-SigLIP2-384__webli (75.67)
Maori nllb-clip-large-siglip__mrl (48.43) None listed
Norwegian nllb-clip-large-siglip__mrl (81.36) ViT-L-16-SigLIP2-384__webli (72.63)
Persian nllb-clip-large-siglip__mrl (79.52) ViT-L-16-SigLIP2-384__webli (74.73)
Polish nllb-clip-large-siglip__mrl (83.49) ViT-L-16-SigLIP2-384__webli (82.03)
Portuguese ViT-SO400M-14-SigLIP2-378__webli (82.12) ViT-L-16-SigLIP2-384__webli (81.39)
Romanian nllb-clip-large-siglip__v1 (89.38) XLM-Roberta-Base-ViT-B-32__laion5b_s13b_b90k (77.92)
Russian ViT-SO400M-16-SigLIP2-384__webli (84.54) ViT-L-16-SigLIP2-384__webli (83.69)
Spanish ViT-SO400M-14-SigLIP2-378__webli (85.47) ViT-L-16-SigLIP2-384__webli (84.81)
Swahili nllb-clip-large-siglip__mrl (69.51) ViT-B-16-SigLIP-i18n-256__webli (21.64)
Swedish nllb-clip-large-siglip__mrl (77.12) ViT-L-16-SigLIP2-384__webli (71.7)
Telugu nllb-clip-large-siglip__mrl (64.32) None listed
Thai nllb-clip-large-siglip__mrl (79.99) XLM-Roberta-Base-ViT-B-32__laion5b_s13b_b90k (66.03)
Turkish nllb-clip-large-siglip__mrl (83.91) ViT-L-16-SigLIP2-384__webli (77.33)
Ukrainian nllb-clip-large-siglip__v1 (83.92) XLM-Roberta-Base-ViT-B-32__laion5b_s13b_b90k (76.31)
Vietnamese ViT-SO400M-16-SigLIP2-384__webli (85.86) ViT-L-16-SigLIP2-384__webli (84.93)

“None listed” means no optimal model for that language fits in about 3 GB.

Model Memory (MiB) Time (ms)
nllb-clip-large-siglip__mrl 4248 75.44
nllb-clip-large-siglip__v1 4226 75.05
XLM-Roberta-Large-ViT-H-14__frozen_laion5b_s13b_b90k 4014 39.14
ViT-SO400M-16-SigLIP2-512__webli 4050 107.67
ViT-SO400M-14-SigLIP2-378__webli 3940 72.25
ViT-SO400M-16-SigLIP2-384__webli 3854 56.57
ViT-L-16-SigLIP2-384__webli 3057 51.7
XLM-Roberta-Base-ViT-B-32__laion5b_s13b_b90k 3030 3.2
ViT-B-16-SigLIP-i18n-256__webli 3029 6.87

Smart search runs in your own machine learning container, and it never moves to Frameleaf Cloud. Hardware acceleration speeds up both indexing and searching; see Hardware acceleration. If your server is short on power, you can run machine learning on another computer; see Remote machine learning.

Smart search models download from Frameleaf’s model mirror, models.frameleaf.cloud, unless you set MACHINE_LEARNING_MODEL_SOURCE_URL or HF_ENDPOINT to your own mirror. If the source doesn’t have the model you choose, loading fails with an error naming the model; it never falls back to another host. See Where models come from.