Image Generation¶
Orb creates an image for an assistant reply when you request one. It uses either an external ComfyUI server or a cloud image provider.
| Backend | You need | Cost |
|---|---|---|
| External ComfyUI | A running server, checkpoint, and imported workflow | Your hardware and electricity |
| Cloud API | A provider account and API key | Provider charges for each image |
Start with ComfyUI Setup or Cloud Image Setup if you have not connected a backend.
Before you start¶
You need:
- A working LLM endpoint in Orb
- A reachable ComfyUI server with a checkpoint and workflow, or a cloud provider connection
The Agent model reads the conversation up to the selected reply and writes the scene part of the image prompt. It also uses the selected style and visible character appearance settings. The image backend receives the final positive and negative prompts, not the conversation, character card, or composition-skill library.
Enable image generation¶
- Open Workflow.
- Select Secondary.
- Turn on Image Generation.
- Select Settings in the Image Generation card.
Connect an image backend¶
Image settings contain Connections and Styles. A connection describes how Orb reaches a backend. A style describes how the image should look and points to a connection. Each style can use a different connection.
ComfyUI¶
- Open Connections and select ComfyUI.
- Enter the server URL, normally
http://127.0.0.1:8188. - Enter a Bearer-token API key if the server requires one.
- Import a workflow and assign it to each style.
- Select a checkpoint for styles whose workflow allows Orb to replace the model.
- Select Test connection, then Save.
The connection test checks each assigned workflow and its required nodes, models, and image files. A ComfyUI connection without an imported workflow cannot render.
Cloud provider¶
- Open Connections → Add connection.
- Choose a provider and enter its API key.
- Select Test connection to load the provider's models.
- Assign the connection to a style.
- Choose a model, resolution, and any supported quality or reference-image options.
- Select Save.
Testing does not generate an image or charge for one. The provider's model, resolution, quality, and reference-image settings belong to the style, so styles can share one connection while using different settings.
Cloud providers may ignore negative prompts, seeds, steps, CFG, samplers, or schedulers. They may also map the requested resolution to a supported size. Render details shows what the provider used and any cost it reported.
Warning
A cloud provider receives the prompt and any reference images you enable. A remote ComfyUI server receives uploaded reference images too. Provider billing and data-retention policies apply.
Import a ComfyUI workflow¶
Orb accepts:
- An API-format JSON workflow
- A ComfyUI output PNG that contains workflow metadata
A regular ComfyUI workflow JSON is not the API format. In ComfyUI, enable the developer options and use Save (API Format) or Export (API).
- Open Image Generation settings.
- Open Imported ComfyUI workflows.
- Select the API-format JSON or metadata PNG.
- Name the workflow.
- Check the positive prompt, negative prompt, seed, image-output, and model slots.
- Optionally map matching
widthandheightinputs. - If the workflow has Load Image nodes, review the reference-image slots.
- Select Confirm slots and add workflow.
- Assign the workflow to a style and choose its checkpoint.
- Set the style's Reference image option when the workflow needs an image.
- Select Test connection, then Save.
Resolution mapping¶
By default, the workflow controls its own output size. Map both width and
height when you want the style's Resolution setting to control a plain latent
node, such as Empty Latent Image.
Leave both unmapped when the workflow gets its size from a reference image,
aspect-ratio node, resolution helper, or a checkpoint-specific setup. Orb only
offers inputs literally named width and height, and both must be mapped for
the style's Resolution field to appear.
Reference images¶
Reference images are off by default. Set the source on a style to use a recent chat image, character references, or both. Character references come from the characters in the scene.1
| Source | Cloud provider | ComfyUI |
|---|---|---|
| Previous image, else character references | One previous image; if none exists, character references up to the provider's capacity | The previous image fills every mapped Load Image node; if none exists, available character references fill the mapped nodes |
| Previous image in the chat | One previous image | The previous image fills every mapped Load Image node |
| Character references | One image per character up to the provider's capacity | One character reference per mapped node; surplus required nodes reuse a character reference |
| Character references and the previous image | Character references and then the previous image, up to the provider's capacity | Character references and then the previous image fill the mapped nodes while space remains |
Character reference images¶
- Open a conversation that contains the character.
- Open Image Generation settings and select This Character Only.
- Choose a PNG, JPEG, or WebP under Reference image.
- Select Save.
The limit is 10 MB. If no reference is saved, Orb uses the character card's avatar.
ComfyUI reference inputs¶
Orb fills the Load Image nodes already present in the workflow. It does not add nodes. With character references enabled, it fills the nodes with separate character references when they are available. If the workflow has more Load Image nodes than available references, the extra nodes reuse one reference so required inputs remain valid.
With Reference image set to Off, ComfyUI uses the filenames stored in the workflow. Those files must exist on the ComfyUI server.
Reference uploads must be PNG, JPEG, or WebP. Cloud references are limited to 4 MB after preparation. A provider may accept fewer references or refuse to use them.
Make an image¶
- Open Workflow → Secondary.
- Choose a style in the Image Generation card.
- Find the assistant reply to visualize.
- Select Visualize reply.
Orb shows whether it is composing the prompt, waiting in a queue, or rendering. Select the image button again to cancel an active request. A submitted ComfyUI job may continue in ComfyUI after Orb cancels it; check that queue before starting another job.
The info button (ⓘ) on an image hides the details and enlarges the image to fill its card. The choice applies to every image and is remembered in this browser. Select ⓘ again to bring the details back.
Download an image¶
Select the download button (⬇) on an image to save it as a PNG.
Orb stores each image as a compressed WebP to save space, and it does not keep a second full-quality copy. For a ComfyUI image, Orb downloads the original PNG from ComfyUI's output folder. That file still contains ComfyUI's workflow, so you can drag it back into ComfyUI.
When the original is unavailable, Orb converts its stored copy to PNG and shows a note. This applies to cloud images, because Orb does not keep the provider's original. It also applies to ComfyUI images when ComfyUI is unreachable, its output file has been deleted, or the file no longer matches.
Variants and rerendering¶
Each image result is a variant of the selected reply.
| Action | Behavior |
|---|---|
| Reroll | Reuses the stored prompt and settings with a new seed when supported. |
| Regenerate | Composes a new prompt using the current style, scene-skill library, and character settings. |
| Rehydrate | Recreates an image whose stored bytes were evicted, using its saved prompt and seed. |
Each cloud action is a new provider request and may be billed. Providers that do not use seeds return a new image for Reroll or Rehydrate.
Changing the selected style before a reroll uses that style's connection and workflow. Orb keeps the stored prompt and records any backend or style change in Render details. Old images keep the resolution used when they were made.
Edit a prompt¶
- Open Render details beside the image. If only the image is showing, select the info button (ⓘ) on it first.
- Select the pencil beside Prompt or Negative.
- Edit the text and click outside the field.
- Select Reroll.
Orb uses your edited text and does not run the prompt-writing step again. Regenerate creates a new prompt and ignores the manual edit.
Character appearance prompts¶
Use This Character Only to add fixed appearance details for a character:
- Positive prompt: visible details such as hair, eyes, body shape, or usual clothes
- Negative prompt: character-specific features to avoid
Do not add a character-count tag such as 1girl; Orb adds the count from the
scene. Style and scene prompts are separate from this character setting.
Both fields take macros. {{char}} means the character
the setting is saved on, so it names the right person in a group chat too.
Styles¶
Orb includes Realistic and Anime styles. Select Add style to make another one. A style contains:
| Setting | Purpose |
|---|---|
| Name | Name shown in the style list |
| Prompt format | Tags, Hybrid, or Prose scene prompts |
| Positive/negative style tags | Visual details to include or avoid |
| Extra instructions | Guidance for the Agent prompt writer |
| Connection | ComfyUI or cloud connection |
| Checkpoint / Workflow | ComfyUI model and workflow |
| Model / Resolution / Quality | Cloud provider settings, where supported |
| Reference images | Images sent to the style's render target |
Tags and Hybrid formats use character-count tags such as 1girl or 1boy.
Prose does not. Match the format to the workflow's text encoder.
The prompt, negative prompt, and extra instructions take macros, resolved for the conversation you generate from. Use them to name a person an edit model receives as a reference image:
Macros resolve when the image is generated, so the prompt in Render details already shows the finished text.
Switching a style's connection keeps its other saved settings. A style can retain both its ComfyUI workflow and cloud model while you switch between them.
Camera and point of view¶
The camera setting is next to the style picker and applies globally:
| Mode | View |
|---|---|
| Auto | A local classifier chooses first-person or third-person from the reply. |
| First-person | Shows the scene through the user's eyes; the user is not drawn. |
| Third-person | Shows the scene from outside and includes the user as a character. |
Orb uses the explicit picker choice first. In Auto mode it uses the classifier, then falls back to third-person when the text is unclear or the classifier is off.
To enable the classifier, open Settings → Local ML, download Auto-POV, and leave it enabled. It runs locally on the CPU.
Composition skills¶
Enable Use scene skills to let Orb select reusable composition guidance before it writes an image prompt. Open Composition skills to enable, add, edit, or remove entries. Each skill has:
- Name: the label shown in settings and Render details.
- When to use: a short applicability summary shown to the selector.
- Instructions: the full geometry, contact, occlusion, framing, crop, or visible-detail guidance shown only to the prompt composer when selected.
Write narrow skills for one composition problem. Explain observable geometry and branching conditions directly, and choose compatible guidance. When to use and Instructions take macros; Name does not, because Orb identifies the skill by it. Orb can select up to four enabled skills and prefers the smallest compatible set. Orb ships a starter library of first-person framings -- front hug, kiss, both back-hug directions, and close-up -- which are ordinary editable entries: change or remove any of them, and a removal stays removed.
When enabled and at least one usable skill exists, selection adds one Agent-model call for each new image or regeneration. A failed or malformed selection does not fail the render; Orb composes without skills and uses its broader subject behavior. Reroll and Rehydrate reuse the stored prompt and skill attribution without selecting again. Regenerate selects again from the current library.
Prompter thinking¶
Enable prompter thinking lets the Agent model reason before it selects scene skills or writes the image prompt. It applies to both calls when selection runs and only to composition when skills are off or no usable skills exist. It can increase token use. Stable thinking settings generally give better prompt-cache reuse.
Both prompt steps use the Agent model's configured Max Tokens as their reply budget. A thinking model spends its reasoning from the same budget, so raise it if prompt steps with thinking on come back empty.
Prompter reference¶
Show the prompter the last image sends the chat's most recent generated image to the Agent model when it writes the prompt, so the outfit, setting, lighting and look carry over between renders even when the story text does not restate them. The prompter composes the final visible instant of the latest assistant reply. It uses the earlier picture only for continuity details the story leaves unchanged, and replaces any pose, action, expression, setting, or framing the story changes. The setting is off by default.
Orb sends the image only when it is on the reply just before the message being visualized, or on a message after that reply. An older picture shows a scene the story has left, so the prompter gets no image then. The message being visualized is skipped, so Regenerate shows the image from the reply before it rather than the render it replaces. Uploads are never sent, because the prompter already sees them in the conversation. When the same picture also goes to the image model as a reference image, the prompter is told so. The earlier image's prompt and negative prompt are not sent. The render's details record which image the prompter saw.
The image travels with the prompt request and the skill-selection call stays text-only, so the cached conversation prefix is unchanged; review turns re-send the same request and reuse it.
Both this setting and Review turns need an Agent model that accepts images. If the provider rejects the image, generation stops with the provider's error instead of continuing without it. A failed review keeps the renders already made and shows the provider error. A server that silently drops images cannot be detected.
With Review turns enabled, each revision's review reason appears only in the live rendering status bar, including while ComfyUI queues the revision. Longer status text automatically scrolls like a billboard; hover or focus the bar to pause it, or hover the status text to read its tooltip. With reduced motion enabled, the bar scrolls manually instead. The reason clears when rendering ends or the next review begins and is not saved in image metadata or review logs. Use the image's variant arrows to browse renders and the generation button's Stop control to end refinement and keep the renders already made.
Troubleshooting¶
| Problem | Try this |
|---|---|
| Image button is missing | Enable the workflow and select an assistant reply without an image. |
| Import a ComfyUI workflow | Import and assign a workflow to the selected style. |
| Choose a checkpoint | Select a checkpoint in each ComfyUI style that needs one. |
| Connection test fails | Check the URL, port, server status, firewall, API key, nodes, and models. |
| ComfyUI completes without an image | Select a valid image-output node during workflow import. |
| Render times out | Check the ComfyUI queue or increase Render timeout from 10 to 900 seconds. |
| Reference image is required | Create or upload an image, or set a character reference and choose a source that includes it. |
| Reference image is too large or unreadable | Use PNG, JPEG, or WebP and a smaller file. |
| Prompt generation fails | Check the LLM endpoint and use a model with tool calling. |
| Prompter reference or Review turns stops with a provider error | Use an Agent model that accepts images, or turn the setting off. |
| Bytes evicted | Select Rehydrate. A cloud backend creates and bills a new image. |
| Provider rejects the key or request | Check the provider dashboard, account credits, model access, and the provider message. |
| Cloud image has the wrong shape | Choose a resolution closer to a ratio supported by that provider. |
| Seed says not used | The provider ignores seeds; this is expected. |
Delete image data¶
Use the delete button in the image header to remove a variant or its variant group. Read Orb's confirmation before continuing.
-
In a group chat, Orb considers characters who have replied in the current round through the selected reply. With scene skills enabled, the selector names the characters actually visible in the shot before Orb chooses reference images. If selection fails, Orb safely falls back to the broader candidate list. If more characters remain than the backend has reference slots, the rest are described in the prompt. Orb supports up to four reference slots per render. ↩