[boxwrench]
BX-77 // Hardware Field Report Filed 2026.08.18

R9700Generation Lab

Measured. Tuned. Preserved.

LTX 2.5, MiniMax H3, and MiniMax Music 3 on AMD's Radeon AI PRO R9700.

BX-77 Boxwrench: a heavy robot knight in burnished gunmetal plate with brass-trimmed pauldrons, a recessed visor slit, and a hexagonal chest core.
Plate 01 — the conditioning anchor

BX-77. This exact image is the reference every conditioned run below is built on — the first frame of run B, the image reference of run C, and the guide frame of run E. Runs A and D never see it, which is the whole comparison.

What we found

Controlled experiments

776×

Model load

931 s → 1.2 s

H3 Qwen3-VL model-load pathology removed with --disable-mmap.

41%

Less wall time

43.78 s → 25.69 s

LTX short-workload result after selecting the 256-token Gemma floor.

24.9%

Faster sampler

29.51 s → 22.15 s

H3 paired Qwen-residency A/B.

These are separate controlled experiments with different boundaries. They are not additive.

The benchmark sequence

2026.08.18 · gfx1201

Four production workflows, executed back to back in one long-lived ComfyUI process. Every wall time below is the measured time for that run in this sequence — not a best-of, not a cold-start figure carried over from the optimization campaign. This is the timing record; the video from these runs was short by design and has been retired in favour of the full-length renders below.

RunWorkflowWallOutput
01LTX 2.511.48 s768×448 · 41f · 8 steps
02H3 T2V131.94 s608×352 · 39f · 20 steps
03H3 I2V49.16 s608×352 · 39f · 20 steps
04Music 328.48 s15 s audio
05H3 R2V81.47 s608×352 · 39f · 20 steps
06LTX return23.62 s768×448 · 41f · 8 steps

326.15 s total · four workflows · successful return to LTX

Field notes — provenance and compatibility issues

Everything above is a run that worked. These are the ones that didn't, kept because a showcase that only records successes is not evidence.

Saved workflows drifted from the live node schema

Two of the golden workflow files predate schema changes in the MiniMax H3 nodes. Both passed /prompt validation and then failed at execution, because ComfyUI validates against the static input table and only resolves the real signature later.

  • H3 image-to-video — the saved graph sends image; the live node wants first_frame. One failed job, then a clean run.
  • H3 reference-to-video — ref_images is a V3 Autogrow input, so the prompt JSON key has to be the dotted dynamic path ref_images.ref_image_0, not the bare slot name and not the legacy ref_image. Six failed attempts before the key format was traced through comfy_api/latest/_io.py. audio_vae and ref_image_size also had to be supplied.

Both were fixed in showcase copies of the workflows. The files under production/ were not modified.

LTX 2.5 — the last latent frame decodes to noise

The 864×480 LTX image-to-video showcase render was clean through frame 160, transitional at 161, and pure noise from frame 162 to the end — measured as high-frequency energy against a blurred copy of each frame: 5.17 at frame 160, 13.72 at 161, then a flat 48.6 for every frame after.

The first fix attempt was LTXVCropGuides, the guide-crop node that was missing between the sampler and the decode. That corrected the frame count from 177 back to the requested 169 and did not touch the noise, which established that the extra frames and the noise had separate causes. Since 161 is 8×20+1 — itself a valid point on LTX's frame grid — and the corruption begins exactly one latent frame later, this is the final latent group decoding degenerately, not a conditioning fault.

The delivered file is the 161 clean frames, re-encoded. No frames were interpolated, time-resampled or synthesised, and the delivered duration is 6.708 s rather than 7.042 s.

Showcase renders were queued against the live node signatures

The same drift that broke the H3 workflows also applies to Music 3: the live MiniMaxMusic3TextEncode takes caption and max_duration, where the saved graph sends prompt and max_audio_frames, and EmptyMiniMaxMusic3LatentAudio requires an explicit batch_size. Every showcase render was queued from a copy under showcase/workflows/; nothing under production/ was modified, and the production preflight still verifies clean.

VAEDecodeAudio tensor-shape mismatch

For certain latent shapes the audio decode raises a size mismatch at dimension 2 and does so deterministically — an identical retry with a fully warm cache fails identically. Runs 01 and 06 ship video-only for this reason. Not a hardware fault, not intermittent.

A run that succeeded and was still wrong

The first attempt at run 01 executed cleanly and produced a valid file of entirely the wrong subject: the production workflow's placeholder prompt text was never overridden for that queued job. It was caught by looking at the extracted frames rather than trusting the status code. The bad file is archived rather than deleted, and run 01 was requeued with the correct prompt.

Session shape

The six runs executed in one long-lived ComfyUI process across one mid-session restart (around 14:58 local, between run 04 and runs 05/06). That restart is why the in-memory /history API could confirm runs 03, 05, and 06 directly and the earlier runs had to be reconciled against on-disk sidecars and file timestamps.

Peak VRAM is recorded as NA for every run except 02. It was not sampled for the manually queued runs, and it is not estimated here.

Full record: final generation acceptance · session.json · session TSV

The showcase renders

864×480 · 24 fps

Below are the presentation renders: full-length demonstrations at 864×480, separated from the short benchmark sequence above.

Why these cost so much more

Roughly seven seconds at 864×480 is about 8.7× the latent volume of a sub-two-second clip at 608×352 — and it does not cost 8.7× more. Attention is quadratic in sequence length, so per-step sampling cost rose about 18.7× between the two sizes. That is the honest reason the benchmark sequence above uses small clips.

ModeConditioningWallResult
T2VText only465.94 sScene lands; identity does not
I2VFirst frame513.89 sStarts exact; identity drifts
R2VImage + voice reference561.80 sIdentity retained
↓ Text only

A

MiniMax H3 — Text to Video

Machine Cathedral

465.94sWall

Machine Cathedral
Poster frame 130, extracted from the delivered file. open source file →

864×480 · 7.29 s · 24 fps · 20 steps · 465.94 s wall

Nothing conditions this run but the prompt text. The architecture, the light and the staging all land as written — and the character is not Boxwrench, and does not stay the same character between shots. That is the honest ceiling of text-only conditioning, and it is the reason the next three runs exist.

Prompt & parameters

Prompt as submitted

integrated_multimodal_description: [Shot 1 | 0.0s-2.2s] A vast machine cathedral stretches into darkness, built from blackened steel ribs, precision-machined vaults, cable conduits, dormant server towers, and monumental archival mechanisms. At the center stands Boxwrench, a robot knight and machine paladin with worn metallic armor, a distinctive mechanical head, layered plated surfaces, visible seams, machined joints, fasteners, and subtle surface wear. Cold cyan diagnostic light traces the edges of the armor while a distant amber furnace glow burns low in the background. The camera begins low and wide, then slowly dollies forward toward Boxwrench. Dust hangs in the air, faint indicator lights wake in sequence along the walls, and the scene feels ancient, intelligent, solemn, and precise. [Shot 2 | 2.2s-4.8s] The camera transitions into a slow three-quarter arc around Boxwrench, revealing more of the immense archive behind him. Concentric holographic signal rings appear far in the distance and travel through the chamber toward him like a message moving backward through time. As the signal approaches, old relays activate, dim control panels flicker to life, and subtle reflections move across scratched metal surfaces, bolts, and engraved armor details. Boxwrench turns his head with calm deliberation, as if receiving the transmission. The motion remains controlled and physically believable, with no exaggerated gestures. [Shot 3 | 4.8s-7.0s] The camera eases closer to a medium shot. Boxwrench raises one armored hand and a compact cold-white point of synthetic light forms above his palm. The light is small but intensely bright, reflecting across the armor and nearby machinery. The distant signal rings dissolve into the surrounding haze. The chamber remains quiet and immense, emphasizing the contrast between cosmic-scale decay and one carefully preserved machine signal. End on Boxwrench holding the light steadily, composed and authoritative. overall_soundscape: Deep machine-room ambience, distant cooling fans, soft structural resonance, occasional relay clicks, one restrained hydraulic movement, faint electrical activation, and a subtle rising transmission tone as the holographic signal approaches. non_diegetic_music: Very restrained dark cinematic ambience with a low sustained tonal bed, distant metallic harmonics, and minimal percussive pulse. Keep it subtle and secondary to the environment.

864×480 · 175 frames · 24 fps · 7.292 s20 steps · res_multistep · simpleseed 811202620.72 s/step · text only, no image or reference

↓ + first frame

B

MiniMax H3 — Image to Video

Signal

513.89sWall

Signal
Poster frame 60, extracted from the delivered file. open source file →

864×480 · 7.29 s · 24 fps · 20 steps · 513.89 s wall

The same weights, now handed the anchor as frame zero. Identity is exact at the start and drifts smoothly from about 2.5 s — measured as mean absolute difference against the anchor, 3.05 rising to 38.46 by the final frame. The silhouette, palette and chest core survive; the helmet does not. Conditioning the first frame buys a beginning, not a whole clip.

Prompt & parameters

Prompt as submitted

For the target video, at 0.00 seconds into the target video, <Picture 1> is fully referenced. integrated_multimodal_description: [Shot 1 | 0.0s-2.4s] Begin exactly from <Picture 1>, preserving Boxwrench's recognizable robot-knight identity, head shape, armor geometry, materials, proportions, mechanical structure, and overall composition. The image should come to life gradually rather than changing abruptly. The camera performs a very slow, steady push inward. Tiny indicator lights begin to flicker subtly. Reflected light shifts across worn metallic surfaces, seams, bolts, machined edges, and layered armor plates. Dust particles drift in the air. The environment remains dark, solemn, and technical. [Shot 2 | 2.4s-4.8s] Boxwrench slowly turns his head as if detecting a distant transmission. A few restrained geometric telemetry elements and cold diagnostic traces appear behind him, staying secondary to the character and never obscuring the silhouette. Small environmental lights wake in sequence. The camera continues a gentle push with a slight three-quarter drift, keeping the motion controlled, stable, and cinematic. Fine surface detail must remain visible throughout: scratches, brushed metal, panel seams, fasteners, and subtle wear. [Shot 3 | 4.8s-7.0s] Boxwrench makes one small deliberate movement with an armored hand, as though acknowledging or receiving a signal. A compact cold-white synthetic light or faint signal glow appears near the hand or reflected across the armor, but remains restrained and elegant. End on a composed medium-close framing with Boxwrench visually stable, authoritative, and intact. The moment should feel precise, intelligent, and quietly mythic. overall_soundscape: Low machine-room ambience, distant ventilation, faint relay clicks, subtle servo movement during the head turn and hand motion, and a quiet high-frequency signal emerging near the end. non_diegetic_music: Sparse dark ambient underscore with a restrained low pulse and subtle metallic harmonic texture. Keep it secondary to the environmental sound. Important constraints: Preserve identity aggressively. Do not redesign the character. No extra limbs, no armor mutations, no face drift, no sudden camera moves, no exaggerated action, no chaotic effects, no random scene change, and no loss of fine material detail. Motion should remain physically plausible, premium, and controlled.

864×480 · 175 frames · 24 fps · 7.292 s20 steps · res_multistep · simpleseed 811202622.21 s/step · first_frame: BX-77 anchor

↓ + image and voice reference

C

MiniMax H3 — Reference to Video

Observatory

561.80sWall

Observatory
Poster frame 95, extracted from the delivered file. open source file →

864×480 · 7.29 s · 24 fps · 20 steps · 561.80 s wall

The strongest of the three. A true reference run — the anchor as <Picture 1> and a twelve-second voice sample as <Audio 1> — and the only one where the helmet, visor, gorget and hexagonal core are still correct in the final frame. The same difference metric stays flat here rather than climbing, because the reference rides through every sampling step instead of only seeding the first one.

Prompt & parameters

Prompt as submitted

For the target video, <Picture 1> is the authoritative visual reference for the character. <Audio 1> is the authoritative reference for the speaking voice: match its timbre, its synthetic metallic resonance, and its measured, unhurried cadence. integrated_multimodal_description: Shot 1, 0.00s to 2.40s. Wide establishing shot inside a high cold observatory carved into dark rock at the end of time. The character from <Picture 1> stands in profile at the centre of a vast circular chamber, small against the architecture, facing a tall aperture open to deep space. Burnished gunmetal plate armour with warm brass trim, a heavy rounded knight close-helm with a deeply recessed horizontal visor slit showing a faint cold red internal optic trace, tall reinforced gorget collar, tiered pauldrons with brass rims, hexagonal mechanical core centred on the chestplate. The light is cold: pale blue starlight falling through the aperture across the floor, with a single low amber instrument lamp behind the character providing the only warmth. Fine dust drifts slowly through the starlight. The camera is locked and still. No movement except the dust and a slow, almost imperceptible pulse in the chest core. Shot 2, 2.40s to 4.80s. Slow, controlled dolly forward to a medium shot as the character turns its head toward the camera with mechanical deliberation. Behind it, along the curved chamber wall, faint engraved concentric rings and thin geometric telemetry traces illuminate one after another in cold cyan, travelling outward and away from the character, reading as a transmission being sent rather than received. The traces stay behind the subject and never cross the silhouette. The character begins to speak, calm and level, in the voice of <Audio 1>: "This is Boxwrench." Shot 3, 4.80s to 7.00s. Composed medium-close shot, the camera settling to a stop. The character raises one armoured gauntlet slowly and holds it steady, palm open, a single compact point of cold white light resting above it. The cyan traces on the wall fade to darkness one by one behind it until only the starlight and the amber lamp remain. The character finishes speaking, unhurried and final, in the voice of <Audio 1>: "The signal holds. Proceed with precision." The final frame holds still on the character and the point of light. overall_soundscape: A vast, cold, quiet interior. Very low sub-bass room tone with a long stone reverb tail. Occasional faint metallic settling and the soft mechanical servo articulation of armour as the character turns and raises its hand. A distant, sparse electronic tick from the instrument lamp. No wind, no crowd, no machinery clatter. The speaking voice sits close and clear in the centre of the field, dry against the large reverberant space behind it, with the synthetic metallic character of <Audio 1> preserved exactly. non_diegetic_music: Extremely restrained. A single sustained low cello or bowed metal drone, entering under shot 2 and holding without swelling. One quiet high harmonic tone appears when the point of light is revealed in shot 3. No percussion, no rhythm, no build, no cinematic hit. The music must never rise above the voice. Important constraints: Preserve identity aggressively. Do not redesign the character. The helmet, visor slit, gorget, pauldrons, brass trim and hexagonal chest core must remain exactly as in <Picture 1> for the full duration. No extra limbs, no armour mutations, no face drift, no helmet opening, no exposed human face, no sudden camera moves, no exaggerated action, no chaotic effects, no random scene change, no loss of fine material detail. No text, no numbers, no user interface, no readouts, and no captions anywhere in the frame. Motion should remain physically plausible, premium, and controlled throughout.

864×480 · 175 frames · 24 fps · 7.292 s20 steps · res_multistep · simpleseed 811202624.12 s/step · ref: anchor + voice, ref_image_size match

ModeConditioningWallResult
T2VText only56.04 sScene lands; generic knight
I2VAnchor frame70.90 sBoxwrench anchored
V2VSource video59.37 sStructure retained; restyled
↓ H3 → LTX, text only

D

LTX 2.5 — Text to Video

The Artifact Forge

56.04sWall

The Artifact Forge
Poster frame 95, extracted from the delivered file. open source file →

864×480 · 6.71 s · 24 fps · 8 steps · 56.04 s wall

LTX with nothing but text, for the same reason run A exists: it is the baseline the next two runs are measured against. The forge lands — hard warm key against a cold shaft, embers, haze, carved stone — and the knight is a plausible knight rather than this one. No brass trim, no gorget geometry, no hexagonal core. Same ceiling as run A, a different model reaching it.

Prompt & parameters

Prompt as submitted

Inside a vast underground forge at the end of time, a towering armoured robot knight works at an anvil of black stone. Burnished gunmetal plate armour with warm brushed brass trim, a rounded medieval close-helm with a deep horizontal visor slit lit by a faint red internal glow, tall reinforced gorget collar, tiered pauldrons with brass rims, heavy segmented gauntlets. It raises a heavy hammer and brings it down on a glowing white-hot ingot; a burst of orange sparks scatters across the frame and dies in the dark air. Molten metal light pulses from a channel running along the floor, throwing hard moving shadows up the carved stone walls. Cold blue light falls from a high shaft far above, cutting through drifting smoke and forge haze. The camera pushes in slowly and steadily from a low three-quarter angle, then settles. Heavy industrial atmosphere, deep shadow, volumetric light shafts, floating embers, fine dust. Cinematic, high detail, film grain, shallow depth of field, warm orange against cold blue, physically plausible motion, no camera shake.

Negative prompt

text, letters, numbers, watermark, signature, user interface, subtitles, captions, human face, exposed skin, extra limbs, extra arms, deformed armour, chaotic motion, fast camera movement, strobing, neon signs, cyberpunk city, crowds, lens flare spam, jitter, stutter, blurry, low quality, oversaturated

864×480 · 161 frames · 24 fps · 6.708 s8 steps · res_multistep · simpleseed 8112026distilled · text only, no image or reference

↓ Same LTX, + anchor frame

E

LTX 2.5 — Image to Video

Signal Vault

70.90sWall

Signal Vault
Poster frame 90, extracted from the delivered file. open source file →

864×480 · 6.71 s · 24 fps · 8 steps · 70.90 s wall

The cost comparison. Same resolution, near-identical duration, guided from the same anchor — and 70.90 s against the 561.80 s the H3 reference run cost. An eight-step distilled model against a twenty-step one, with no audio branch. The last latent frame decoded to noise and was trimmed; the detail is in the field notes.

Prompt & parameters

Prompt as submitted

A heavy armoured robot knight stands motionless at the centre of a deep underground vault, facing a tall sealed archive door. Burnished gunmetal steel plate armour with warm brushed brass trim along every bevel and rim, a rounded medieval close-helm with a deeply recessed horizontal visor slit lit by a faint cold red internal glow, a tall reinforced gorget collar, tiered pauldrons with brass edges, and a hexagonal mechanical core set into the centre of the chestplate. The vault is carved from dark stone and lined with rows of dormant metal storage racks receding into shadow. A single low amber lamp burns behind the figure, throwing a long soft shadow forward across a dusty floor; thin cold cyan light bleeds from seams in the sealed door ahead. Fine dust drifts slowly through both light sources. The camera pushes forward very slowly and steadily toward the figure from behind and slightly to the left, revealing the scale of the vault. The knight raises one gauntleted hand and rests it flat against the sealed door, and holds it there. The chest core pulses once, softly. Cinematic, high detail, restrained, solemn, film grain, shallow depth of field, warm amber against cold cyan, no camera shake.

Negative prompt

text, letters, numbers, watermark, user interface, subtitles, captions, human face, exposed skin, extra limbs, deformed armour, chaotic motion, fast camera movement, neon signs, cyberpunk city, crowds, lens flare spam, jitter, stutter, blurry, low quality

864×480 · 161 frames · 24 fps · 6.708 s8 steps · res_multistep · simpleseed 8112026distilled · guide: BX-77 anchor at frame 0

↓ Restyle, not regenerate

F

LTX 2.5 — Video to Video

Observatory, restyled

59.37sWall

Observatory, restyled
Poster frame 95, extracted from the delivered file. open source file →

864×480 · 6.71 s · 24 fps · 8 steps · 59.37 s wall

No empty latent and no guide frame — run C's finished video is VAE-encoded straight into the starting latent, and the sampler only partially renoises it. That is the difference between restyling something and generating it: 59.37 s here against the 561.80 s the source cost.

Prompt, the denoise trade-off & parameters

Prompt as submitted

Cel-shaded anime style. A heavy armoured robot knight stands in a vast cold observatory chamber at the end of time, facing a tall aperture open onto a deep starfield. Hand-drawn Japanese animation aesthetic: bold clean ink linework, flat cel shading with hard-edged shadow shapes, limited colour banding instead of smooth gradients, painted matte background art, subtle paper grain. Burnished gunmetal armour rendered as flat grey planes with crisp white specular highlights and warm brass trim picked out in solid gold. A rounded knight close-helm with a deep horizontal visor slit and a single glowing red optic. Cold blue starlight from the aperture against one warm amber lamp, drawn as distinct pools of flat colour rather than photographic falloff. Thin cyan geometric telemetry rings on the chamber wall drawn as clean vector arcs. Cinematic anime key-art composition, high production value, studio quality, dramatic but restrained, film grain.

Negative prompt

photorealistic, photograph, 3d render, cgi, realistic skin texture, smooth photographic gradients, live action, text, letters, numbers, watermark, signature, user interface, subtitles, deformed anatomy, extra limbs, blurry, low quality, jitter, stutter, chromatic aberration

Denoise is the whole trade-off

The ask was cel-shaded anime. At the denoise that keeps the source video, you do not get it — you get stylised illustration with anime cues. The fully hand-drawn look is reachable, but only by giving the source up. Both ends were run:

  • denoise 0.6 — shipped above. SHIPPED. Source structure preserved — same camera, same staging, same telemetry rings, same character pose through the clip. Style genuinely changed: flat planar shading, heavier outlines, solid gold trim, graphic vector arcs, painted starfield. Reads as stylised key art rather than strict cel anime.
  • denoise 0.85 — not shipped. NOT SHIPPED as the primary. Style is fully 2D — bold ink outlines, flat cel colour, painted background — but the source is essentially gone: different composition, different camera, character re-staged front-on and simplified. At this denoise it stops being a restyle and becomes a fresh generation that happens to share a prompt.

Stronger style transfer and structure retention are directly opposed here, and eight distilled steps leaves little room to have both. The 0.85 take is kept because it shows that clearly: open the 0.85 version →

864×480 · 161 frames · 24 fps · 6.708 s8 steps · res_multistep · simpleseed 8112026 · denoise 0.6distilled · source: run C, VAE-encoded

↓ Past the window

—

Continuation

Two clips, joined at the frame

14.55sLongest

Both models are bounded — H3 by its frame grid, LTX by the sequence length that stays coherent. Neither is bounded at the story. So the last frame of a finished clip becomes the first frame of the next one, and the two are joined.

The joins measure 3.56 and 3.03 out of 255 — inside codec noise. The continuation reproduces the frame it was handed almost exactly, so these read as one continuous take rather than a cut. The duplicated frame is dropped before joining; there is no crossfade, no interpolation and no time remapping.

Observatory, extended
14.546 s · 349 frames · seam at 7.29 s. The light lifts from the gauntlet and leaves through the aperture. Audio carries across the seam. open source file →
The Artifact Forge, extended
13.375 s · 321 frames · seam at 6.71 s. The ingot comes off the anvil and is carried to the molten channel. LTX, so no audio. open source file →
ModelOutputGeneration timePer second of video
H314.55 s, with audio18m 11s75.0 s
LTX13.38 s, no audio2m 06s9.5 s

H3 took 18m 11s to produce 14.55 s. LTX took 2m 06s to produce 13.38 s. That is an 8.6× difference in total wall for roughly comparable output length, or 7.9× once normalised per second of video.

seed 8112027 · 1091.00 s and 126.42 s exact

↓ Video → Music

G

MiniMax Music 3

Proceed with Precision

185.99sGeneration wall

90.02sAudio

MP3, LAME V0, 90.02 s, 44 kHz stereo. Open source file →

Waveform of the 90 second render: dense, near-constant high-energy transients throughout, consistent with blast-beat drumming and distorted guitars.

Waveform, rendered from the delivered MP3. Integrated loudness -14.16 LUFS, range 5.5 LU, true peak +0.67 dBTP — hotter than the earlier render and clipping at the inter-sample level; not re-mastered for this page.

“Proceed with precision, through the fire and the noise…” Chorus — lyric excerpt

Ninety seconds this time, three times the previous render's length. 185.99 s wall doesn't scale linearly with duration alone — sampling steps went up 20 → 30 and tiled VAE decode was switched on for the longer render, both adding cost on top of the 3× audio. That puts it at roughly a third of an H3 video run's wall time, but it's now slower than any single LTX video run above — no longer the cheapest thing on this page.

Prompt, lyrics & parameters

Caption (style prompt)

Brutal 1990s old-school death metal. Down-tuned buzzsaw guitars, tremolo-picked riffs and chugging palm-muted breakdowns, blast beats and double-kick drums, thick distorted bass. Deep guttural death growl vocals, low and monstrous, shouted gang backing vocals on the hook. Raw analog production, dry mix, no modern polish, no clean singing, no melodic choruses. Fast and heavy, around 180 BPM, minor key, with a half-time breakdown.

Lyrics

[Intro] [Solo lyrical cello, rising D minor melody] [Distorted guitars and galloping drums enter] [Verse] I walked the crooked road beneath a broken sky Where every mile was measured by the things I left behind The stones were cut by doubt, the bridges burned by fear But every scar became a mark that brought the future near [Pre-Chorus] No hand can turn the river No king can still the tide The past becomes a compass When purpose is your guide [Chorus] Proceed with precision Through the fire and the noise Proceed with precision Let the future hear your voice From the crooked path behind us To the line we now envision We rise, we build, we burn Proceed with precision [Verse] The old machines are sleeping in the dust beneath the rain Their iron hearts remember every failure, every name But in the pulse of wire and light, a clearer compass turns A thousand lost equations become lessons as we learn [Pre-Chorus] The wheel is still in motion The dark gives way to dawn The map is drawn in action The road is never gone [Chorus] Proceed with precision Through the fire and the noise Proceed with precision Let the future hear your voice From the crooked path behind us To the line we now envision We rise, we build, we burn Proceed with precision [Solo] [Cello and harmonized lead guitar exchange the main melody] [Galloping drums and rhythm guitars continue] [Bridge] I am not the shadow Of the road beneath my feet I am every broken answer Made relentless, made complete [Instrumental] [Solo cello carries the melody] [Brief rising transition into final chorus] [Final Chorus] Proceed with precision Through the fire and the noise Proceed with precision Let the future hear your voice From the crooked path behind us To the line we now envision We rise, we build, we burn Proceed with precision [Outro] [Solo cello reprises the opening melody] [D major resolution]
Spectrogram of the 90 second render showing dense broadband energy across the low and mid range, consistent with distorted guitars and blast-beat drumming.

seed 101226310236506030 steps · euler · simplecfg 1.7 · top-k 50max duration 90.0 s · tiled decode

41m 54s of measured generation time 2513.51 s exact · nine renders · one R9700

Every figure above is read out of showcase/metadata-v2, which also records what each run got wrong.

We found some weird stuff.

If yours behaves badly

Model load takes fifteen minutes?

Check mmap before blaming your SSD. --disable-mmap.

H3 encoder 931 s → 1.2 s

H3 gets slow after a new prompt?

Check whether Qwen is still resident before touching the sampler.

Paired sampling A/B 29.51 s → 22.15 s

Short LTX prompt takes twenty seconds?

You may be processing a 1,024-token minimum sequence you never asked for.

Selected: LTX_GEMMA_MIN_LENGTH=256

Music looks GPU-idle for twenty seconds?

The work is in the autoregressive conditioner, not the diffusion transformer.

Measured: ~19.8 s AR loop vs ~3.6 s DiT

Full troubleshooting field guide →

Known-good state

Preflight verified
Workloads
LTX 2.5text-to-video, 22B distilled DiT, INT8-ConvRot
MiniMax H3T2V / I2V / R2V, FP8-scaled, Qwen3-VL 32B encoder
MiniMax Music 3INT8-ConvRot DiT, discrete audio codec
ComfyUI
v0.33.27cee3ceb1a35503172e0dfb8dbdbdedee2aba8aa
Torch
2.9.1+rocm7.2.1
ROCm / HIP
7.2.53211-e1a6bc5663
Triton
3.5.1+rocm7.2.1
GPU
AMD Radeon AI PRO R9700gfx1201 · 31.86 GiB
Model hashes and provenance

SHA-256, truncated. Full digests, sizes, and workflow fingerprints live in the manifest and are re-checked by the preflight script on every run.

FileRoleSHA-256 (16)
ltx-2.5-22b-distilled-transformer-comfy-int8-convrotLTX 2.5 DiTc4279eeff115cbea
gemma4-12b-with-proj-ltx-2.5-comfy-int8-convrotLTX 2.5 text encoder09a89e084de1a149
ltx-2.5-video-vae-bf16LTX 2.5 video VAE847e14ca7f3355de
minimax_h3_fl2va_pruned_fp8_scaledH3 first/last-frame transformer12944c1f7791637e
minimax_h3_ref2va_pruned_fp8_scaledH3 reference transformerf86f2f79ebd2d76e
qwen3vl_32b_minimax_h3_fp8H3 Qwen3-VL encoder5e1127e510a30099
minimax_h3_video_vae_fp16H3 video VAE7c1f131492e7edda
minimax_h3_audio_vae_fp32H3 audio VAE8e505d95dd1561d4
minimax_music3_dit_int8_convrotMusic 3 DiTd6b959633e69899f
minimax_music3_text_encoder_pruned_int8_convrotMusic 3 lyric encoder010b7416d2336a08
minimax_music3_davMusic 3 audio codec2a32155b769be014

production/manifest.json

Make your R9700 do this.

Start here

Verify the machine matches the recorded state before anything else. The preflight checks hardware, runtime versions, applied patches, model hashes, and workflow fingerprints in one pass.

python3 scripts/production-preflight.py

End of field report

BX-77 standing alone in profile at the centre of the dark observatory chamber, small against the architecture, lit by cold starlight from a tall aperture and a single low amber lamp.
Plate 02 — from run C

BX-77 // End of field report

Known-good state preserved.

Proceed with precision.