When Human Preference Becomes Part of the Training Loop
Rapidata proposes replacing a learned image reward model with live pairwise judgments during online alignment. The interesting engineering question is not only whether people can respond quickly, but how to keep their signal reliable as the model changes.

Live human feedback can reduce dependence on a stale reward proxy, but it needs explicit criteria, uncertainty tracking, governance, and held-out evaluation to remain trustworthy.
A community article published October 2 describes a system for inserting live human preferences into online post-training for visual generative models. Rapidata presents its Flows API as a substitute for the learned reward model in a Flow-GRPO-style loop: generate several images for one prompt, gather pairwise judgments, turn those comparisons into a score for each image, normalize the scores within the group and use the resulting advantages for an update. The article says groups can be evaluated asynchronously and reports that its service can return ratings under short deadlines. Those performance and capacity statements come from Rapidata and are not independently verified here.
The proposal addresses a genuine weakness of proxy rewards. A fixed reward model represents preferences observed before the current policy produced its newest outputs. As optimization moves the generator into unfamiliar regions, a proxy can become unreliable or reward artifacts that exploit its learned shortcuts. Asking people about current outputs shortens that distance, but it does not eliminate measurement problems. It moves the critical design work from choosing a reward model to operating a human measurement system.
The feedback target must be explicit
A comparison such as “Which image looks better?” is easy to answer but underspecified. One annotator may reward visual polish, another prompt fidelity and another realism. If the intended product needs legible text or faithful brand colors, a broad attractiveness question can push training away from the actual requirement.
Write the evaluation rule before collecting votes. A useful instruction names the primary criterion, defines what to do when both candidates fail and clarifies whether the prompt should be visible. Complex goals may be safer as separate signals rather than one overloaded choice. For example, collect prompt adherence and visual defects independently, then decide how the training objective should combine them. That makes tradeoffs visible instead of hiding them inside annotator intuition.
The same care applies to prompts. If a batch overrepresents portraits, improvements in that slice do not establish broader model quality. Track the prompt mixture by content, language, safety sensitivity and difficulty. A changing mixture can otherwise look like a change in model quality even when the model is unchanged.
Latency is only one systems constraint
Live feedback puts a variable-rate human process on the critical path of a predictable compute job. Asynchronous collection can overlap image generation with rating, but the trainer still needs rules for late, incomplete or contradictory responses. A timeout should not silently turn a difficult example into a weak preference, because difficult examples may be precisely where the signal is most valuable.
Design the loop with explicit states: submitted, sufficiently rated, expired and excluded. Record the number of judgments behind every score, the response-time distribution and the agreement level. If the system proceeds with partial results, mark them and test whether weighting by confidence changes the update. The right threshold is workload-specific; the key is that it is chosen and audited rather than inherited from an API default.
Throughput planning also needs a cost envelope. Estimate generated groups per training step, comparisons required per group, acceptable wait time and expected retry or rejection rates. Run a small pilot across busy and quiet periods before reserving expensive accelerators. A system that is fast on average may still create costly stalls at the tail.
Pairwise scores are measurements, not truth
Rapidata says Flows decompose a group ranking into pairwise comparisons and fit a Bradley-Terry model, producing Elo-style scores. This is a convenient bridge from discrete votes to group-relative advantages, but the output remains an estimate shaped by which pairs were shown and who evaluated them. Close scores should not be treated as a precise ordering without uncertainty information.
Before training, replay the scoring pipeline on controlled groups. Include obvious wins, near ties, duplicated images and deliberately corrupted outputs. Duplicates can reveal position or presentation bias; repeated groups can reveal instability. Break results down by prompt slice and interface condition. If judgments depend strongly on image order, screen size or whether context is shown, changing the interface mid-run changes the reward function.
Human review also introduces governance requirements. Confirm that raters are allowed to see the generated material, especially when prompts or outputs may contain personal, disturbing or confidential content. Define filtering, escalation and retention policies before collection begins. Faster feedback is not a reason to bypass those controls.
Keep an independent evaluation outside the loop
Online preferences are used to change the model, so they cannot also provide an unbiased final verdict. Hold out prompts and evaluators from training, then compare checkpoints on criteria chosen in advance. Include automated checks where they directly measure requirements, such as duplicate rate or text legibility, but do not collapse every result into one headline score.
Monitor diversity as well as preference wins. Repeatedly selecting the most broadly appealing output can narrow style, composition or subject coverage. Compare per-slice win rates, defect rates and diversity indicators against the starting checkpoint. Preserve losing examples and disagreement cases; they are often more diagnostic than the average score.
A sensible rollout starts with shadow mode: collect live judgments without updating weights, inspect reliability and estimate latency and cost. Next, run short bounded updates with frequent checkpoints and a rollback criterion. Only then extend the training horizon. Live human feedback can make the signal more current, but its value depends on treating annotation, statistics and operations as parts of the training system rather than as an interchangeable API call.
Source: Live Human Feedback in the Training Loop: Aligning Diffusion Models with Real People ↗. How we write


