I asked for 4.5:1 contrast and the model returned 2.59:1
The contrast requirement was in the prompt, stated as a number. One model returned an accent colour at 2.59:1 against the background it had just chosen, and on another run declared the scheme was light while giving a near-black background.
TL;DR · THE FIX
A prompt is a request, not a constraint. Asked for body text at 7:1 and an accent at 4.5:1, a small model returned an accent at 2.59:1 and, on another description, a JSON object declaring scheme light with a #333333 background, a contradiction inside a single response. The fix is not a better prompt: compute the contrast in the code that renders the palette, repair anything that fails by moving lightness until it passes, and show the visitor what was corrected. Validate generated values against the property you needed, in code, every time.
The symptom
A small feature: describe the look you want in a sentence, get a colour palette back, see it applied. The prompt asked for JSON and stated the accessibility floor as numbers:
Body text must reach a contrast ratio of at least 7:1 against the background. The accent colour must reach at least 4.5:1 against the background.
What came back, from Llama 3.1 8B on Workers AI, was well-formed JSON with every field present and plausible hex values in all of them. It parsed, it rendered, and at a glance it looked like a palette. The accent measured 2.59:1 against the background in the same object.
On a different description, the same model returned this:
{ "scheme": "light", "bg": "#333333", "fg": "#f5f5f5", "accent": "#7ec8e3" }
scheme: "light" next to a background of #333333 contradicts itself inside one JSON object, and it would have driven the light-mode branch of the UI while painting a near-black background.
Nothing errored. The response was valid JSON with correct keys and correct types, and every check I had was a check on shape.
What I did first, and why it was the wrong instinct
The first reaction is to argue with the model: restate the requirement, add an example and a counter-example, tell it to compute the ratio before answering, ask it to include the computed ratio in the output. Some of that helps and none of it settles anything. The prompt is a request, and the model is not running a contrast formula. It produces text that looks like the text that would follow that request, and a hex triple that looks right is exactly what comes out of that process. Asking it to report the ratio too gets you a number that looks like a ratio, sitting next to colours it does not describe.
I did compare models. Haiku 4.5 cleared both bars on all three test descriptions, and it read the references properly: “swiss watch advert” produced a Swiss red on white monochrome rather than the generic tech cyan the smaller model reached for. The difference is real, and building on it would still have been the mistake. A model that passes three test descriptions has passed three test descriptions. The next one is a fresh sample from a distribution that includes 2.59:1, and shipping it to visitors lets that sample decide whether your text is readable.
The fix
Compute the thing you needed, in the code that renders it. Contrast ratio is a short, exactly specified function:
const srgbToLinear = (c) => {
const v = c / 255;
return v <= 0.03928 ? v / 12.92 : Math.pow((v + 0.055) / 1.055, 2.4);
};
function luminance(hex) {
const n = parseInt(hex.slice(1), 16);
const [r, g, b] = [(n >> 16) & 255, (n >> 8) & 255, n & 255];
return 0.2126 * srgbToLinear(r) + 0.7152 * srgbToLinear(g) + 0.0722 * srgbToLinear(b);
}
export function contrast(a, b) {
const [x, y] = [luminance(a), luminance(b)].sort((p, q) => q - p);
return (x + 0.05) / (y + 0.05);
}
Then repair anything that misses, rather than rejecting the whole response and asking again. Rejection costs a round trip and returns another sample from the same distribution. Repair is deterministic and instant: keep the hue and saturation the model chose, which is the part it is good at, and move the lightness until the ratio passes.
function enforce(fg, bg, target) {
if (contrast(fg, bg) >= target) return { hex: fg, repaired: false };
const darken = luminance(bg) > 0.5; // dark text on a light background
let { h, s, l } = hexToHsl(fg);
for (let i = 0; i < 100; i++) {
l = Math.max(0, Math.min(100, l + (darken ? -1 : 1)));
const candidate = hslToHex(h, s, l);
if (contrast(candidate, bg) >= target) return { hex: candidate, repaired: true };
if (l === 0 || l === 100) break;
}
return { hex: darken ? "#000000" : "#ffffff", repaired: true };
}
One degree of lightness at a time keeps the result as close to the model’s intent as the requirement allows, and the fallback to black or white guarantees termination, because pure black or pure white always clears 4.5:1 against anything that is not almost itself.
The scheme contradiction gets the same treatment. A field that can be derived should be derived:
const scheme = luminance(bg) > 0.5 ? "light" : "dark"; // derived, not read
If a value in a response can be computed from another value in the same response, compute it. A model can contradict itself within one object and arithmetic cannot.
Then tell the user. The repair pass names every correction in the interface:
Accent adjusted for contrast (2.59:1 to 4.6:1). Scheme set to dark from the background colour.
Without that line the visitor asked for something, got something visibly different, and has no way to know why, which reads as a broken feature rather than a working guardrail.
The lesson
A generated value has to be re-validated against the property you needed, by the code that consumes it. The prompt expresses an intention and has no enforcement power, so treating it as though it does means your accessibility floor is a sentence in a string literal that nothing checks.
For any generated field, ask whether it is a preference or a constraint. Hue, mood, and which reference the description evokes are preferences, and a model is the right tool for them. A contrast ratio, a sum that must balance, a date that must fall in a range, an identifier that must exist are constraints, and each of those can be verified in a few lines of ordinary code. Verify the constraints, keep the preferences, and repair rather than retry, because retrying asks the same process for a different sample.
Discussion
Powered by GitHub. Sign in to leave a comment.