This is a bench for learning about prompt engineering when using text-to-image and image-to-image diffusion models. It is currently in DRAFT form... I'm not tracking breaking changes.
This project uses uv to manage the Python version, the virtual environment, and dependencies. If you don't have uv and direnv installed, start with that:
# install uv (see the uv docs for other install methods)
brew install uv direnv
# uv reads the pinned interpreter from .tool-versions (3.12.9)
# and will install it for you on the first sync
uv python install 3.12.9Then you can create the virtual environment and install packages with direnv:
# configure your ENVVARS
# (open the .envrc.local file and make necessary changes)
cp .envrc.local-example .envrc.local
# create the venv (.venv), run `uv sync`, and enable ENVVARS
direnv allow .
# If you don't want to use direnv, read the .envrc file and
# run those commands (mainly `uv sync`) using your preferred methodDependencies are declared in pyproject.toml and locked in uv.lock. To add a
dependency, use uv add, which updates both files:
uv add scipyuv sync (run automatically by .envrc) installs everything from the lockfile.
To upgrade locked versions, use uv lock --upgrade followed by uv sync.
This project pins diffusers to the HuggingFace main branch (via
[tool.uv.sources] in pyproject.toml) so the latest conditioning models and
detectors — including TencentARC/t2i-adapter-lineart-sdxl-1.0 — are available.
python3 -m src [-h] [--prompt PROMPT] [--negative_prompt NEGATIVE_PROMPT]
[--width WIDTH] [--height HEIGHT] [--count COUNT]
[--seeds SEEDS] [--models MODELS]
[--custom_latents CUSTOM_LATENTS] [--steps STEPS]
[--output_path OUTPUT_PATH]
[--output_path_template OUTPUT_PATH_TEMPLATE]
[--input_paths INPUT_PATHS] [--device_type DEVICE_TYPE]
[--refinement_mode REFINEMENT_MODE] [--copyright COPYRIGHT]
options:
-h, --help show this help message and exit
--prompt PROMPT, -p PROMPT
A pipe-delimited list of descriptions of what you would
like to render
--negative_prompt NEGATIVE_PROMPT, -n NEGATIVE_PROMPT
A pipe-delimited list of descriptions of what you would NOT like to render (default=None)
--width WIDTH, -x WIDTH
The width of the output image (default=896)
--height HEIGHT, -y HEIGHT
The height of the output image (default=640)
--count COUNT, -c COUNT
The number of images to produce with each given model
(default=1)
--seeds SEEDS, -s SEEDS
A comma-separated list of PRNGs to use when generating an
image to produce more predictable results and explore an
idea
--models MODELS, -m MODELS
A comma-separated list of HuggingFace Models to to use. By
default, dreamlike-art/dreamlike-photoreal-2.0 is used to
generate an image followed by 3, sequential steps of
refinement using stabilityai/stable-diffusion-xl-
refiner-1.0
--custom_latents CUSTOM_LATENTS, -l CUSTOM_LATENTS
A comma-separated list of booleans: whether or not to use
custom latents (seeds) for each model
--steps STEPS, -r STEPS
A comma-separated list of the number inference steps to use
with each model (length must match the length of -m)
--output_path OUTPUT_PATH, -o OUTPUT_PATH
A path to a folder where images will be saved (can be
relative)
--output_path_template OUTPUT_PATH_TEMPLATE, -t OUTPUT_PATH_TEMPLATE
A template for naming the files
(default=":path/:count_idx-:type-:model_idx.png")
--input_paths INPUT_PATHS, -i INPUT_PATHS
A comma-separated list of paths to images that will be
refined or upscaled by the given models
--device_type DEVICE_TYPE, -d DEVICE_TYPE
The type of device the pipes will be fed to for processing
(default="cuda" if cuda is supported, else "mps" if apple
M1/M2, else "cpu")
--refinement_mode REFINEMENT_MODE
one of: "sequence" (each pass is fed into the next pass),
"first_to_many" (each pass is fed the first item
generated), or "in_to_many" (each pass is fed the value of
-i/--input_paths)
--copyright COPYRIGHT
Set who should be listed as the copyright owner of the
images that are createdGeneration is resource-hungry and can make the machine hard to use for anything
else. run.sh wraps python3 -m src to run it as a lower-priority job: it still
gets most of the CPU when the machine is idle, but yields to interactive work,
and it leaves a couple of cores free for the system. All arguments are passed
straight through:
./run.sh --prompt "a cute, tabby cat" \
--models "dreamlike-art/dreamlike-photoreal-2.0" \
--steps "10"
# run it detached and log output:
./run.sh --prompt "a cute, tabby cat" --models "..." --steps "10" > run.log 2>&1 &Tunables (override via environment):
DIFFUSION_BENCH_LEAVE_FREE— cores to leave free for the system (default:2)DIFFUSION_BENCH_NICE— scheduling priority; higher yields more (default:15)PYTORCH_MPS_HIGH_WATERMARK_RATIO— Apple Silicon only. Caps how much unified memory MPS may use so the GPU/UI stays responsive and the machine doesn't swap. Unset by default because too low will OOM large models (e.g. HiDream); try0.7for the smaller SD / dreamlike models.
Transform an .mp4 with a prompt. Pass a video to -i/--input_paths.
Coherent (recommended) — AnimateDiff. Use -m guoyww/animatediff-v1-5-2.
A MotionAdapter keeps the output consistent frame-to-frame. It's SD1.5-based and
its attention cost scales with resolution squared, so the processing size is
auto-capped to 512px (aspect-preserving); raise it with
DIFFUSION_BENCH_MAX_SIZE=768 ... if you have the memory. First run downloads
~2.5GB of weights.
For video, -x/-y act as a maximum bounding box, not an exact size: the
clip's own aspect ratio (and orientation, including rotated phone footage) is
always preserved. Omit them and the source is fit inside a 512px box.
Clips of any length are processed in overlapping windows (crossfaded at the
seams), so memory is bounded by one window regardless of duration. --window_size
(default 16) is the memory/quality knob — larger windows are more coherent but
use more memory; --overlap (default 4, must be ≤ half the window) controls the
blend. --max_frames just trims how much of the clip is used.
./run.sh --prompt "analog style, a tabby cat" \
--models "guoyww/animatediff-v1-5-2" \
--input_paths "clip.mp4" \
-x 512 -y 512 --strength 0.6 --guidance_scale 8.5 \
--fps 12 --max_frames 32Naive fallback. --naive refines each frame independently with any
img2img-capable model — cheap, but flickers (no temporal coherence).
./run.sh --prompt "analog style, a tabby cat" \
--models "timbrooks/instruct-pix2pix" \
--input_paths "clip.mp4" \
--naive --strength 0.6 --guidance_scale 8.5 \
--fps 12 --max_frames 48Video options:
--naive— route video through the per-frame path (needs an img2img model)--strength— change amount,0.0–1.0(higher = further from the source)--guidance_scale/-g— prompt adherence--fps— output frame rate (subsamples when lower than the source)--max_frames— cap frames processed--window_size/--overlap— window/blend size for the coherent path (unused by--naive)--controlnet—off|lineart|depthstructure conditioning (coherent path)
Output is written as .mp4 with a <name>.mp4.json sidecar recording the
prompt, model, and parameters (mp4 has no EXIF equivalent).
The coherent path generates at ~512px, so a separate upscaler restores
resolution. It's a standalone step (python3 -m src.upscale, or ./upscale.sh),
not part of the generation pipeline, and works on both images and video.
# upscale a generated clip back to its source resolution
./upscale.sh -i images/clip-animatediff.mp4 --to 1920x1080
# or a plain multiplier, on an image or video
./upscale.sh -i frame.png --upscaler realesrgan --scale 4Upscale options:
--upscaler/-u—realesrgan(default; sharp, deterministic) orlanczos(soft, free fallback)--to— target size as WIDTHxHEIGHT (e.g.1920x1080), aspect preserved--scale— multiplier (e.g.2); used when--tois omitted--fps/--max_frames— video only--weights— override the Real-ESRGAN weights (e.g. an anime variant); the defaultRealESRGAN_x4plus.pthis auto-downloaded to~/.cache/diffusion-bench
Note: upscaling synthesizes plausible detail — it can't recover what wasn't
generated. Real-ESRGAN is per-frame but deterministic, so on the already-coherent
output it adds little flicker; video-native backends (RealBasicVSR/BasicVSR++)
and diffusion upscalers are planned (see docs/plans/upscale.md).
models:
- wavymulder/Analog-Diffusion
- NOTE: you have to use "analog style" in the prompt for this to take effect
- To use a EulerAncestralDiscreteScheduler append "/EulerA" to the model name: "wavymulder/Analog-Diffusion/EulerA"
- Deci/DeciDiffusion-v1-0
- dreamlike-art/dreamlike-photoreal-2.0
- prompthero/openjourney
- stabilityai/stable-diffusion-2-1
- stabilityai/stable-diffusion-2-depth
- stabilityai/stable-diffusion-x4-upscaler
- stabilityai/stable-diffusion-xl-base-1.0
- stabilityai/stable-diffusion-xl-refiner-1.0
- TencentARC/t2i-adapter-lineart-sdxl-1.0
- timbrooks/instruct-pix2pix
- minimaxir/sdxl-wrong-lora
- NOTE: you have to use "wrong" as a negative prompt for this to take effect
Given the same prompt, model, and number of inference steps, generate 3 images:
python3 -m src --prompt "a cute, tabby cat" \
--models "dreamlike-art/dreamlike-photoreal-2.0" \
--steps "10" \
--count 3
# =>
# loading pipeline: dreamlike-art/dreamlike-photoreal-2.0
# Loading pipeline components...: 100%|███████| 5/5
# [00:01<00:00, 4.06it/s]
# prompt: ['a cute, tabby cat']
# negative_prompt: []
# width: 896
# height: 640
# count: 3
# seeds: []
# model_ids: ['dreamlike-art/dreamlike-photoreal-2.0']
# steps: [10]
# input_paths: []
# paths: [['images/out-01-GENERATOR-01.png',
# 'images/out-02-GENERATOR-01.png',
# 'images/out-03-GENERATOR-01.png']]
# device: mps
#
# model_id: dreamlike-art/dreamlike-photoreal-2.0
# prompt: a cute, tabby cat
# kwargs: {'negative_prompt': None, 'num_inference_steps': 10,
# 'width': 896, 'height': 640}
# seed: 4042405154439330027
#
# 100%|████████████████████████████████████| 10/10
# [00:35<00:00, 3.52s/it]
# ...Given the same prompt, model, and number of inference steps, generate 3 images each, for 2 different models: "dreamlike-art/dreamlike-photoreal-2.0" and "stabilityai/stable-diffusion-xl-base-1.0":
python3 -m src --prompt "a cute, tabby cat" \
--models "dreamlike-art/dreamlike-photoreal-2.0, stabilityai/stable-diffusion-xl-base-1.0" \
--steps "10, 10" \
--count 3Given the same prompt, model and seed, generate images using a different number of inference steps for each image:
python3 -m src --prompt "a cute, tabby cat" \
--models "dreamlike-art/dreamlike-photoreal-2.0, dreamlike-art/dreamlike-photoreal-2.0, dreamlike-art/dreamlike-photoreal-2.0, dreamlike-art/dreamlike-photoreal-2.0" \
--steps "10, 15, 20, 25" \
--seed "1175243130925179488, 1175243130925179488, 1175243130925179488, 1175243130925179488" \
--count 1Given a base image with a blurry, mangled face:
python3 -m src --prompt "a cute, tabby cat" \
--models "dreamlike-art/dreamlike-photoreal-2.0" \
--seeds "12621049111909025765" \
--steps "10"Generate 4 images using a different negative prompt for each image ("blurry", "mangled face", "deformed face", "blurry, mangled face, deformed face"):
python3 -m src --prompt "a cute, tabby cat" \
--negative_prompt "blurry | mangled face | deformed face | blurry, mangled face, deformed face" \
--models "dreamlike-art/dreamlike-photoreal-2.0, dreamlike-art/dreamlike-photoreal-2.0, dreamlike-art/dreamlike-photoreal-2.0, dreamlike-art/dreamlike-photoreal-2.0" \
--seeds "12621049111909025765, 12621049111909025765, 12621049111909025765, 12621049111909025765" \
--steps "10, 10, 10, 10"Given the same negative prompt, model, seed, and number of inference steps, generate 4 images with different prompts for each image (a dog, a cat, a bear, a pig):
python3 -m src --prompt "a dog | a cat | a bear | a pig" \
--negative_prompt "outside" \
--models "dreamlike-art/dreamlike-photoreal-2.0, dreamlike-art/dreamlike-photoreal-2.0, dreamlike-art/dreamlike-photoreal-2.0, dreamlike-art/dreamlike-photoreal-2.0" \
--steps "10, 10, 10, 10" \
--seed "1175243130925179488, 1175243130925179488, 1175243130925179488, 1175243130925179488" \
--count 1NOTE the bear has the most outside context. Is this because the model is trained with pictures of domesticated animals both inside and outside, and with wild animals that are primarily outside?
Given a prompt:
- generate an image
- refine the image generated in step 1
- refine the image refined in step 2
python3 -m src --prompt "a cute, tabby cat" \
--models "dreamlike-art/dreamlike-photoreal-2.0, stabilityai/stable-diffusion-xl-refiner-1.0, stabilityai/stable-diffusion-xl-refiner-1.0" \
--steps "10, 50, 50" \
--refinement_mode "sequence"Given an input image:
- load the image from disk
- refine the image loaded in step 1
- refine the image refined in step 2
python3 -m src --prompt "a cute, tabby cat" \
--input_paths "images/01-generator-01.png" \
--models "stabilityai/stable-diffusion-xl-refiner-1.0, stabilityai/stable-diffusion-xl-refiner-1.0" \
--steps "50, 50" \
--refinement_mode "sequence"Given a prompt:
- generate an image
- refine the image generated in step 1
- refine the image refined in step 2
python3 -m src --prompt "a cute, tabby cat | replace the cat with a lion" \
--models "dreamlike-art/dreamlike-photoreal-2.0, timbrooks/instruct-pix2pix" \
--steps "10, 10" \
--refinement_mode "sequence"
--seed "16340881476748913736"Make sure to read the licenses for any model you use with this. The licenses vary from model to model. For instance dreamlike-art/dreamlike-photoreal-2.0 uses an adaptation of Creative Commons and limits corporate usage.
This bench should not be used to intentionally create or disseminate images that create hostile or alienating environments for people. This includes generating images that people would foreseeably find disturbing, distressing, or offensive; or content that propagates historical or current stereotypes.








