EyeSayYour account
← Your account

EyeSay image API: tags, sorting, search vectors, alt text

Everything the page does, from your own code — or from an assistant you connect.

Access

Sign in and open your account: the key on it is what your code sends, as the header X-Auth-Token. Keep it secret; if it leaks, replace it there and the old one stops working at once.

An assistant or other app does not need your key. It connects with OAuth: it finds everything it needs at /.well-known/oauth-authorization-server, you approve it on a page here, and it sends Authorization: Bearer …. You can disconnect it from your account page.

Requests

Every call is a POST with a JSON body to /v1/<endpoint> on this site. A photo is a base64 string or a data URL. It can also be upload:<batch>/<n>: a photo you put on the upload page, usable by your account for thirty minutes. A request may be up to 32 MB, and a photo up to 16 million pixels.

curl https://<this site>/v1/tags \
  -H "X-Auth-Token: $KEY" -H "Content-Type: application/json" \
  -d '{"image": "data:image/jpeg;base64,...", "k": 8}'

Endpoints

/v1/tags — {"image", "k"?}
Tags for a photo: {"tags": [{"tag", "p"}], "min_score"}, k from 1 to 20. A tag whose p is below min_score is a suggestion. Add "with_embedding": true to get the photo’s search vector in the same answer, counted as one photo. Send {"image", "vocab": [...]} instead to score your own list of words.
/v1/tags/aggregate — {"image", "metadata"?, "k"?}
Tags for one photo, merged with facts you already know about it (for now {"is_screenshot": true}), each tag saying where it came from.
/v1/classify — {"image", "labels": [...], "k"?}
Sorts the photo into your own labels, up to 64 of them, with a score for each.
/v1/embed — one of {"image"}, {"images": [...]}, {"text"}, {"texts": [...]}
Vectors for search, up to 64 inputs a call. Photos and text land in the same space, so a text vector finds photos. Keep the space value with what you store: vectors from different spaces must not be mixed.
/v1/rerank — {"query", "documents": [...], "top_k"?}
Re-orders up to 128 candidates (photos or text) against a query, best first.
/v1/ask — {"image", "text"?}
No text: a one-sentence description (caption or alt text). A yes/no question: the probability of yes. Any other question: a short answer. Text that is not a question (a statement to match against the photo) is refused with 400 and costs nothing.
/v1/analyze — {"image", "text"?, "labels"?}
A description, tags and a search vector in one call; with labels, their scores too, and with text, how well the photo matches it.

Every answer carries model_version, so you can tell when results may shift.

The same endpoints for tools and code generators: OpenAPI spec.

Describe a photo (caption / alt text)

Send a photo to /v1/ask with no text and you get one plain sentence about it — the same sentence Describe writes on the home page. Use it as a caption, or as the alt text of the image on a web page.

curl https://<this site>/v1/ask \
  -H "X-Auth-Token: $KEY" -H "Content-Type: application/json" \
  -d '{"image": "data:image/jpeg;base64,..."}'
{"mode": "caption", "answer": "a brown dog running on a beach",
 "confidence": 0.71, "model_version": "..."}

The sentence is in answer; mode is caption when a description was written. It may start in lower case and end without a full stop: for alt text, add both, as the Alt text line under Describe on the home page does. Rely on the fields shown here.

A description is 9 units, so €0.90 per 1,000 photos; a free account’s month covers 333. It takes one photo per request, within the limits above (up to 32 MB, 16 million pixels, 60 requests a minute); for many photos, send them one per call, a few at a time.

The sentence is written by AI and can be wrong, or miss what matters in the photo. Read each one before you publish it as alt text.

Assistants (MCP)

An assistant that speaks the Model Context Protocol connects to /mcp on this site and signs in with OAuth as above. Its tools list your uploaded photos, tag them, sort them into your labels, find the ones matching a description, and describe or answer questions about one photo. Each photo a tool handles is one request, counted like any other. A tag the model is not sure of comes marked "suggestion": true.

Use with Claude on the web, desktop and mobile

Add EyeSay once in Claude on the web or in Claude Desktop. It is then available in the Claude apps for iOS and Android as well.

  1. In Claude, open Customize → Connectors, click +, and choose Add custom connector.
  2. Name it EyeSay, enter https://eyesay.app/mcp as the URL, and click Add. Leave Advanced settings empty.
  3. Find EyeSay in the list and click Connect. A page on eyesay.app opens: sign in with your EyeSay account and allow access.
  4. In a chat, click + in the lower left, open Connectors and turn EyeSay on for that conversation.

On Claude’s Free plan you can add one custom connector. On a Team or Enterprise plan, an owner first adds it under Organization settings → Connectors.

To give Claude your photos, upload them at https://eyesay.app/upload, signed in with the same EyeSay account, then ask Claude to use them, for example “tag the photos I just uploaded”. Uploads are kept in memory for 30 minutes at most. Claude sends EyeSay up to 20 photos at a time and goes on with the rest.

Use with Claude Code

For a folder of photos on your computer, the photos should never pass through the assistant: that is slow and costs far more. There are two ways.

With the connector (nothing to set up). Add it once: claude mcp add --transport http eyesay https://eyesay.app/mcp, then sign in with /mcp. Ask “use EyeSay to tag the photos in ~/Pictures/trip”; Claude Code may ask you to confirm once before the photos are sent. The assistant calls prepare_upload, which starts a photo job and returns a one-line command; the command sends each photo in the folder straight here (shrunk to 896 px first on a Mac), where it is tagged as it arrives, and then prints the result: the counts, the most frequent tags and a link to a CSV of every photo (job_summary gives it again later). The assistant can pass job:<id> to classify_photos or search_photos to sort or search the whole folder without sending it again; photos close to none of your labels go to Unsorted. To put them into folders, it runs the organize command the sort returns (Claude Code may ask you to confirm once): each photo is copied to EyeSay/<label>/ inside your folder, your originals are never moved, renamed or deleted, and deleting that EyeSay folder undoes it. Results are kept 24 hours; running the command again sends only what is missing. A photo job is not limited per minute: it takes 4 photos at a time per account (2 at a time across the site); the 60-a-minute limit below is for the API.

If auto mode blocks the command. In the default permission mode Claude Code asks you once and you click to allow it. In auto mode, its safety check may block the command instead, because the upload and organize commands download a script and run it. You can then choose either of these:

With your own key. Put your key in an environment variable (export EYESAY_API_KEY=…, never paste it into the chat) and let the assistant write a loop that posts each photo to /v1/tags, writes the full answers to a local file and prints only a summary. Send JPEG or PNG (convert HEIC first, for example with sips -s format jpeg on a Mac) no larger than 896 px on the long side, at most a few at a time; on a 429 with Retry-After, wait and retry; without it, the month’s photos are used up. A short version of all this for assistants is at /llms.txt.

Use with ChatGPT

With EyeSay on in a ChatGPT conversation, attach the photos to your message and ask, for example “tag these photos” or “sort these into food and receipts”. ChatGPT hands EyeSay up to 20 photos per call and goes on with the rest; our server fetches each attached photo from OpenAI’s file server, answers it in memory and keeps nothing (see privacy). You can also upload photos at https://eyesay.app/upload, signed in with the same EyeSay account, and ask ChatGPT to use the photos you just uploaded.

Use with Codex

Add EyeSay once with codex mcp add eyesay --url https://eyesay.app/mcp, then sign in with codex mcp login eyesay: a page on eyesay.app opens, where you sign in with your EyeSay account and allow access. Folders work as in Claude Code above: ask “use EyeSay to tag the photos in ~/Pictures/trip”, and Codex runs the upload command prepare_upload returns, and later the organize command. Those commands send the photos over the network, so Codex asks you once to approve network access for them.

Export tags as XMP sidecars

For a photo job (a folder sent by an assistant, above), the tags can be exported as XMP sidecar files: one small .xmp file per photo, named like the photo with the extension replaced (IMG_0042.jpg gets IMG_0042.xmp), that Lightroom, Bridge and other catalogues which read XMP sidecars take as keywords. Only text is involved: the photos are not sent again and not changed.

Write them next to your photos. Add --xmp after the folder in the upload command that prepare_upload returns (--xmp-suggestions adds the suggestions):

curl -fsSL https://eyesay.app/intake.sh | sh -s -- "<job URL>" "<token>" "<FOLDER>" --xmp

On Windows, add --xmp (or -Xmp) after the folder in the PowerShell command. After the tagging summary it prints how many sidecars were written. An .xmp that already exists is never overwritten (it is counted and left alone), no folder is created, and the photos are never modified. Without the flag nothing is written.

Or download them. job_summary returns results_xmp, a signed link (valid for 15 minutes, like the CSV and JSON links) to /intake/<job>/xmp.zip: a zip of <folder path>/<name>.xmp with the paths relative to the folder that was sent. Unzip it into that folder so each sidecar lands beside its photo. The zip is built in memory when you fetch it and nothing extra is stored. A file that was sent under two names (a copy in another folder) gets a sidecar for each name; if two photos share a sidecar name (IMG_1.jpg and IMG_1.png), their keywords are joined in one file.

In Lightroom. Sidecars next to raw files are read when the photos are imported. For JPEG and other formats that Lightroom does not read sidecars for on import, import the photos, select them and choose Metadata → Read Metadata from File. Reading metadata from a file takes over what the file holds for that photo, so save any changes you made in the catalogue first.

Allowance

Photos are counted in units, priced by what is done with each photo: a search vector (embed, rerank) is 1 unit, tags or a sort into your labels (tags, classify) 3 units, a description or an answer (caption, ask, analyze) 9 units. A unit is €0.0001, so a thousand photos cost €0.10 as vectors, €0.30 tagged and €0.90 described. Text sent to embed is free.

A free account has 3,000 units each calendar month — 1,000 tagged photos, or 3,000 search vectors, or 333 descriptions — at up to 60 requests a minute; while it has bought units left, up to 600. When the month’s units run out, calls are paid from units bought on your account page, which do not expire. Your account page shows what is left. Uploading is free; a photo counts when a call uses it. A photo sent into a photo job counts as one tagged photo (3 units), when it arrives, for its tags and search vector together; the same photo sent again does not count. An account takes at most 50,000 image vectors a day (UTC); write to support if you need more.

Units are sold in packs, paid once: 100,000 units (about 33,000 tagged photos) for €10, 420,000 for €40, 1,650,000 for €150. Checkout is handled by Lemon Squeezy, which shows any sales tax before you pay; refunds follow the terms.

Answers that are not results

400
The body is not what the endpoint takes; the message says what is wrong.
401
No key, a wrong key, or an expired app pass.
403
The endpoint is not part of your account’s plan.
413
The request is too large; send fewer or smaller photos.
429
Too many requests, or the service is busy — wait the number of seconds in Retry-After and retry — or, with no Retry-After, this month’s allowance is used up, which the message says.

GET /health says whether the service is up; HEAD /v1/tags says whether tagging can be served right now. Neither needs a key or counts against anything.

What happens to your photos: privacy. See also the terms.