Skip to content

Set up AI descriptions

This page is for administrators. It covers turning on AI descriptions, choosing a model, tuning the prompt and backfilling an existing library. The settings live in Administration, then Settings, then Machine Learning Settings, then Image descriptions and tags.

  • Your machine learning container is working, ideally with hardware acceleration. See Hardware acceleration.
  • The model you pick is a Qwen2.5-VL or Phi-3.5-vision model, not Florence-2.
  • For names in descriptions, facial recognition is on and you’ve named some people. See People.
  • Videos need nothing extra. Each video’s moment frames are cut the first time it’s described.
Step Do Why now
1 Pick the right model Florence-2 ignores every other setting on this page. Get the model right before you tune anything.
2 Run facial recognition and name your most-photographed people Descriptions only use names you’ve given. Doing this first means the first run already says “Kelly” instead of “Someone”.
3 Preview a description on a few photos and videos Nothing is written, so you see the result before paying for a library-wide run.
4 Tune the prompt Tuning before the big run means you don’t describe everything twice.
5 Turn on identity injection Cheap to change; set it together with the prompt.
6 Re-queue all descriptions One library-wide pass with everything set the way you want.
7 Turn on smart albums Smart albums fill from description tags, so they need descriptions first.
8 Re-evaluate smart albums once descriptions are done A one-time backfill for older photos.

If you’ve already started without following this order, that’s fine: you can re-queue at any time, and smart albums can be re-evaluated one album at a time.

  1. Go to Administration, then Settings, then Machine Learning Settings.
  2. Under Image enrichment hardware, choose Auto-detect, Intel iGPU (OpenVINO) or NVIDIA GPU (CUDA).
  3. In Image descriptions and tags, make sure Generate image descriptions and tags is on.
  4. Choose a Description model. The default, Qwen/Qwen2.5-VL-3B-Instruct, suits most servers.
  5. Open Prompt & Vocabulary, then Identity injection, and make sure Enable identity injection is on.
  6. Save.
  7. Use Preview a description to check the result on a few of your photos and videos.
  8. Under Status & Re-generation, choose Re-queue all image descriptions, check the estimate, and confirm.
  9. Turn on smart albums if you want them.

New uploads are described automatically after their thumbnails are made, using your current settings.

Hardware Model Notes
NVIDIA with 16 GB or more video memory Qwen/Qwen2.5-VL-7B-Instruct Best quality of the common choices
NVIDIA with 6 to 12 GB Qwen/Qwen2.5-VL-3B-Instruct The default; a good balance
Intel integrated graphics Qwen/Qwen2.5-VL-3B-Instruct Runs as llmware/qwen2.5-vl-3b-ov through OpenVINO
CPU or small integrated GPU microsoft/Phi-3.5-vision-instruct Lighter, still follows the prompt
Older Phi build microsoft/Phi-3-vision-128k-instruct Smaller, slightly lower quality
Last resort microsoft/Florence-2-base-ft Captions only

The Description model list shows an estimate of the video memory each model needs:

Model Video memory Notes
Qwen2.5-VL 3B (default) About 6 GB Good at objects and scenes. On OpenVINO it runs as llmware/qwen2.5-vl-3b-ov.
Qwen2.5-VL 7B About 16 GB Better at composition, text in images and subtle scenes. On OpenVINO it runs as llmware/qwen2.5-vl-7b-ov.
Qwen2.5-VL 32B About 64 GB Complex scenes and fine detail. NVIDIA only.
Qwen2.5-VL 72B About 144 GB Best quality in the family; needs several GPUs. NVIDIA only.
Qwen3-VL 30B-A3B About 60 GB Only about 3B parameters active at a time, so it runs near 7B speed. NVIDIA only.
Phi-3.5-vision-instruct About 5 GB Smaller alternative on OpenVINO
Phi-3-vision-128k About 5 GB Older and smaller still
Florence-2-base-ft About 1 GB Captions only; local fallback
Florence-2-large-ft About 3 GB Captions only; local fallback

Custom… lets you type any Hugging Face model ID, but only the Qwen2.5-VL, Qwen3-VL, Phi-3 and Phi-3.5-vision, and Florence-2 families load. Anything else fails with Failed to load model in the worker logs.

A model downloads the first time a job uses it, which can take several minutes.

Fallback model is used on this server when the main model can’t run. Florence-2 is the usual choice because it’s small enough to share a GPU with other models. If the main model fails on a local machine learning server, Frameleaf retries the same request with the fallback. Frameleaf Cloud always uses its chosen model, with no fallback.

On Intel, use the OpenVINO machine learning image and leave the description device on AUTO, which picks the best device and falls back when the GPU isn’t available. On NVIDIA, use the CUDA image and choose NVIDIA GPU (CUDA).

Description models download from huggingface.co unless you set HF_ENDPOINT to your own mirror. See Where models come from. Your photos are processed by your own machine learning container, and Frameleaf sends no telemetry.

Preview a description opens Try an enrichment change:

  1. Choose sample. Pick up to six of your own photos or videos, and where to run them. The destination is always named; nothing falls back to another one.

  2. Compare. The model and prompt you’re editing, saved or not, run on each sample. The current description and the new one appear side by side, with the tags, any names the check removed and, for a video, how many frames it saw. Nothing is written: descriptions, tags, the Locked state, search data and frames stay as they are.

  3. Scope. Choose the stages a background plan should run on the samples:

    • Descriptions and tags, and the Locked-content check, for photos
    • Reusable video frames, the moment search index and moment captions, for videos

    The moment stages need the frames, so ticking one ticks the frames too. Moment captions are never ticked for you: they add one model request per frame, and the dialog says how many.

A plan uses the saved model and prompt, and records them on every result it writes. It runs in the background one item at a time, shows in Activity, and survives closing the browser or restarting the server. You can Pause or Cancel it between items. Each failure is retried once automatically, and Retry starts a new plan over only the items and stages that didn’t finish.

Setting What it does Default
Description style Terse (one or two sentences), Balanced (a short paragraph) or Rich (a detailed description) Balanced
Sentence count target Target number of sentences, from 1 to 6. The model sometimes goes one over. 3
Look for Categories the model should point out when they’re visible brands, signage, screens, documents, uniforms, tools, vehicles, animals, food, landmarks
Custom vocabulary Tag values the model should reuse, spelled exactly as you write them Empty
Custom instructions Your own guidance in full sentences, up to 2,000 characters Empty
Forbidden inferences Things the model must not infer, even when the photo suggests them diagnoses, medication names, procedures, pregnancy, disability
NSFW indicators Explicit terms allowed in descriptions of sensitive photos. Clear the box and save to restore the defaults. A built-in list
Medical indicators Medical terms allowed in descriptions. Clear the box and save to restore the defaults. A built-in list
Identity injection Names recognised people in descriptions On, 5 names, 0.7
Advanced (raw prompt editor) Replaces the whole prompt with your own template Off

List settings take one entry per line.

Library Style Identity injection Add to Look for Add to Custom vocabulary
Family Balanced or Rich On birthday cake, sparklers, candles, prom, graduation, recital, soccer ball, baseball glove golden hour, overcast, dappled light, backyard, park, beach
Documents and receipts Terse Off total amount, store name, line items, payment method, invoice number, due date, expiry date, signature line Keep it short
Travel Balanced On mountain range, beach, hotel lobby, train station, food market, street scene, landmark The landmarks and cities you visit often
Pets Terse or Balanced Either breed, leash, food bowl, harness, fetch, asleep, swimming, kennel Your pets’ breeds, such as golden retriever or tabby

For a documents library, also add account numbers verbatim, full credit card numbers, social security numbers to Forbidden inferences.

For a hobby, teach the model its words through Custom vocabulary: long exposure, leading lines, bokeh for photography, latte art, sourdough, plating for cooking, or gravel, derailleur, peloton for cycling.

Custom instructions is the easiest way to change how the model writes without touching the whole prompt. Treat it like a short brief to a careful assistant: write full sentences, say what to do and when not to, and keep it to a couple of paragraphs. The text is added to the prompt before the output rules, so the model reads it first.

Some examples that work well:

Vehicles

If you can clearly see a car, truck, or motorcycle, identify the make and model in
the description (e.g. "a red Tesla Model 3", not "a red car"). If you are uncertain
about either the make or the model, just say "a red car" rather than guessing.

Sports

When people are clearly playing a sport, name the sport in the description (soccer,
baseball, basketball, tennis, etc.) and mention the visible equipment they're using.
Do not guess the sport from clothing alone, only from visible play, equipment, or
field markings.

Travel

When the photo is clearly outdoors and looks like travel, identify the location type:
beach, mountain, forest, city street, market, train station, airport, museum, cathedral,
temple, ruins, hotel lobby. If a recognizable landmark is visible, name it. Do not
invent landmarks if you are not sure.

Documents and receipts

When the photo is a document, receipt, or screenshot of text, transcribe the most
important visible identifiers: store name, total amount, date, invoice number, and
document type. Do not transcribe full account numbers, full credit card numbers,
social security numbers, or any other sensitive personal data even if visible; refer
to them generically (e.g. "account number redacted").

Food

For food photos: identify the dish by name if recognizable, name the cuisine if you
can tell, and note the plating style (rustic, fine-dining, casual, street-food). Do
not name a dish you are not confident about; say "pasta dish" rather than guessing.

Pets

For pets: identify the breed when you can recognize it. For dogs in particular, note
the activity (sleeping, playing fetch, swimming, on a leash, in a car). Do not guess
a breed from coat color alone.

Tone

Keep descriptions plain and factual. Do not use poetic language ("a tapestry of
colors", "bathed in golden light"). Do not editorialize about the people or events
("a joyful family", "an intimate moment"). Stick to what is visible.

A family library, all in one

Always name every recognized person and avoid generic group terms like "the family"
or "everyone".
If you see a car, name the make and model when recognizable.
If people are playing a sport, name the sport and any visible equipment.
For pets, identify the breed when recognizable.
For travel scenes, name the landmark if you are certain.
Keep descriptions factual and short. Do not use poetic language.
  • They can’t override the safety rules. The sensitive-content and medical handling, and the check that removes invented names, still apply.
  • They can’t add names. A name that isn’t recognised on the photo is still replaced with “Someone”.
  • They can’t change the output format. Use the raw prompt editor for that.
  • They can’t make the model remember anything between photos. Each one is described on its own.
  • They’re ignored while the raw prompt editor is on.
Setting What it does Default
Enable identity injection Passes the named people in each photo to the model On
Max names How many people to name in one photo, from 1 to 20 5
Min face confidence The confidence a face match needs before its name is used, from 0 to 1 0.7

Lower Max names to 1 or 2 for crowd photos where only the main people matter. Raise it to 10 to 15 for teams, weddings and other big groups. If a photo has more named people than Max names, some are left out.

Most people should leave the Advanced raw prompt editor alone. The other settings build the same prompt with safer defaults. Use it only when you need something they can’t express, such as a different output format or sections in another order.

  • When you turn on Enable raw prompt override with an empty box, Raw prompt template fills with the current default template, so you start from something that works. It never overwrites text you’ve already written.
  • Reset to default replaces your template with the default, discarding your edits.
  • Turning the override off goes back to the structured settings. Your template is kept for next time.

You can use these placeholders:

Placeholder Becomes
{names} The recognised people, with the wording that requires the model to use them. Empty when identity injection is off or nobody is recognised.
{schema} The output format the model must return. Required in strict mode.
{vocabulary} Your custom vocabulary
{style_hint} The length and tone cue for the chosen style

Placeholder validation is either Strict, where saving fails without {schema}, or Warn, where it saves and shows a warning. Strict is the default.

Custom instructions aren’t added to a raw template; paste the text where you want it. Frameleaf still adds the video timing note for videos, and extra care for photos flagged as sensitive, around your template.

If the template is wrong, descriptions come out empty or broken. Turn the override off to recover.

Status & Re-generation shows:

Field Shows
Last config change When the description settings last really changed
Pending re-queue scheduled Set when you chose Re-queue later
Total eligible image assets Everything descriptions can process
Already described Items with a description
Pending re-description Items without one
Estimated re-queue time Based on the last 100 finished description jobs

Re-queue all image descriptions opens a window with the counts, the model, the seconds per item and the estimated total time:

  • Cancel closes it.
  • Re-queue later leaves a reminder banner at the top of the settings, so you can make several changes and run once at the end. Re-queue now on the banner starts it.
  • Re-queue now starts straight away. Clicking twice doesn’t queue the work twice.

When descriptions go to Frameleaf Cloud, nothing is queued here; Frameleaf Cloud describes photos in batches instead.

After a restart, the estimate uses 1.5 seconds per item until 100 jobs have finished. That’s about right for Qwen2.5-VL 3B on a typical Intel integrated GPU. On an NVIDIA GPU, expect around 0.5 to 1 second per item with the 3B model.

You can also run Image descriptions and tags, then All, from Administration, then Jobs. Jobs skip items that already have a description unless you force them, and a forced run never adds the AI description: block twice.

Describe video moments, under the Descriptions & tags routing settings, is off by default. When it’s on, every video described automatically also gets its frames captioned straight afterwards, one extra model request per frame. It doesn’t change any description, so turning it on or off never asks for descriptions to be redone. Frames looked at per video beside it shows the sampling policy, six frames, and can’t be changed.

If Detect NSFW images is also on, Frameleaf checks each photo for sensitive content first and passes the result to the description model, so the description stays factual. See Private content.

Say you have about 80,000 photos and 3,000 videos, and one NVIDIA GPU with 8 GB of memory.

  1. Day 1. Turn on facial recognition, run face detection and wait for it to finish. Name your 20 or so most-photographed people; the People page shows them first.
  2. Day 2. Choose Qwen/Qwen2.5-VL-3B-Instruct, which fits comfortably in 8 GB. Set Balanced and 3 sentences, add a few things to Look for and Custom vocabulary, paste the family example into Custom instructions, and set Max names to 8. Save, then preview on a handful of photos and videos.
  3. Day 3. Re-queue all descriptions. At 0.5 to 1 second each, expect 12 to 24 hours. Turn on smart albums, add any extra triggers, and click Re-evaluate all assets when descriptions finish; that takes minutes.
  4. Day 4. Open ten random photos and videos. Are names used? Are your instructions followed? Do smart albums make sense? Do video descriptions use words such as “begins” and “then”? Adjust, rerun a few items from their info panels, and repeat until you’re happy.
  5. After that. New uploads are described automatically. New names appear in new descriptions straight away; rerun older ones if you want them updated.

Descriptions are still generic with identity injection on

Section titled “Descriptions are still generic with identity injection on”
  1. Check the photo has at least one named face in its People list.
  2. Check the model is Qwen or Phi, not Florence.
  3. Check Max names is at least the number of named people in the photo.
  4. Rerun descriptions and tags for the photo from the Image enrichment section of its info panel. Older descriptions keep the wording they were written with.
  1. Check the raw prompt editor is off.
  2. Check you saved.
  3. Rerun some photos. Existing descriptions don’t change by themselves.
  4. Try plainer, more direct wording: “If you see a car, identify the make and model” works better than “try to be more specific about vehicles”.
  5. Check you’re under 2,000 characters; longer text can’t be saved.

A video shows “video-frames-unavailable”

Section titled “A video shows “video-frames-unavailable””

No frames could be cut. The video is very short, longer than three hours, or unreadable. Choose Find moments in its Moments section to try again. If that fails, check that video processing works for other files and that the file isn’t damaged.

Turn the override off and on again; it fills itself only when the box is empty. Or click Reset to default.

It clears when a re-queue starts. If the job failed, check Administration, then Jobs, then click Re-queue now on the banner.

After a restart the estimate uses a default until 100 jobs have finished, then settles over the next 100 or so.

A real name was replaced with “Someone”

Section titled “A real name was replaced with “Someone””

The person isn’t named on that photo. Name them in its People list and rerun it, or edit the description by hand. There’s no separate list of allowed names.