How to use
- Download the model once, then add images.
- Choose Short alt text (up to 125 characters) or Detailed description, optionally a context prefix, and press Describe.
- Edit each description or mark decorative images, then copy the text, the <img> HTML or a CSV.
Worked example
The sample city illustration (lit tower blocks under a full moon) gets the short alt text “A graphic of a building with a moon in the background.” (54 characters).
Supported formats and limits
| Input | JPEG, PNG, WebP |
|---|---|
| Output | Alt text (copy), <img> HTML, CSV for batches |
| Limits | Up to 30 images per batch. The first use downloads the Florence-2-base-ft model from Hugging Face (about 230 MB), cached afterwards. It runs on the CPU (WebAssembly): expect several seconds per image. |
| Engine | Florence-2-base-ft (MIT, ONNX by onnx-community) via transformers.js in a Web Worker (WebAssembly) |
Limitations
- The model can misidentify objects, counts and colours, and it does not read text in images reliably; always review.
- It does not identify people; it may guess gender from appearance unless you choose neutral words.
- Alt text should describe the image's purpose on your page, which the model cannot know.
Questions
Can I publish the alt text without checking it?
No. The descriptions are machine-generated by Florence-2-base-ft, a small vision model. It can misidentify objects, counts and colours and does not read text in images reliably. It also cannot know why the image is on your page, which good alt text should reflect.
What is downloaded, and is my image uploaded?
Florence-2-base-ft (MIT licence, about 230 MB) is downloaded once from Hugging Face and cached. Images are processed on your device and are not uploaded. It runs on the CPU, so expect several seconds per image.
How does it handle people?
It does not identify who someone is, but it does guess gender from appearance. Turn on Neutral words for people to replace words such as man or woman with person, people or child.
Privacy
On-device model. Processing runs in this browser. The open-source model files are downloaded once from the model host (Hugging Face) and cached; your content is not uploaded.
- Hugging Face: Only after you press Download model: requests for the model files (Florence-2-base-ft, about 230 MB) from huggingface.co and its CDN, which see your IP address. Your text and files are never sent.
See the privacy policy for how toolsdocks handles data.