Compliance
Sensory testing
"It tastes better" is not data. Sensory testing turns taste into numbers you can track across batches, compare against a gold standard, and defend to a buyer. None of it needs a laboratory β it needs discipline: the same questions, the same conditions, the same scale, every time.
Which test for which question
| Question | Test | Panel | Output |
|---|---|---|---|
| Is this batch different from last month's? | Triangle or paired-difference | Trained or semi-trained, 8β15 | Different / not different (statistical) |
| How does the new dryer change the product? | Descriptive profiling | Trained, 6β10 | Attribute scores per sample |
| Do customers like it enough to buy again? | Hedonic (9-point) test | Target consumers, 30β100 | Acceptance score, purchase intent |
| Has it degraded in storage? | Difference vs fresh reference + attribute scores | Trained, 6β10 | Shelf-life endpoint |
| Which of three recipes goes forward? | Ranked or hedonic | Semi-trained or consumers | Preference order |
Panel conditions β the boring part that decides everything
- Neutral environment Quiet, odour-free room; no cooking smells, perfume or ambient music. Natural light or neutral white for colour judgement.
- Coded samples Three-digit random codes, never names; randomised presentation order (different order per assessor).
- Controlled portions Equal sample sizes at room temperature β warm samples taste sweeter and softer; define the temperature and keep it.
- Cleansers Plain water and unsalted crackers between samples; 30β60 s pauses.
- Spit-out and ethics Provide spit cups for trained panels; only taste food that passed the safety checks β sensory panels never sample suspect product.
- No discussion Assessors record independently before any conversation; discussion contaminates scores.
Difference tests: is it actually different?
Triangle test
Each assessor receives three samples β two identical, one different β and picks the odd one. With 8β12 assessors, the count of correct answers is compared to a standard table (for n = 9, 6 correct answers is significant at p < 0.05). If the panel can't reliably pick the odd sample, your "improvement" is in your head β which is exactly the kind of thing worth knowing before you change the process.
Paired comparison
Two samples, one question: "which is sweeter?" Fast, decisive, ideal for A/B testing a process change (e.g. dipped vs undipped colour). Count preference; significance tables decide.
For an in-house panel, use 9 assessors and the triangle test; the standard tables for 9β12 assessors are easy to find and give a genuinely defensible answer. Run it blind even when you "know" the answer β especially then.
Descriptive profiling: what changed
A trained panel scores defined attributes on a line or 0β10 scale. For dried food, a good generic profile covers appearance, texture, aroma and taste:
| Attribute | Anchor 0 | Anchor 10 |
|---|---|---|
| Colour intensity | Bleached / grey | Vivid, true to type |
| Uniformity | Wildly variable pieces | Identical pieces |
| Firmness (bend) | Crisp / snaps | Very soft / pliable |
| Chewiness | Dissolves instantly | Long, resistant chew |
| Acidity | None | Sharp, face-puckering |
| Sweetness | None | Intense |
| Aroma intensity | Flat, cardboard | Intense, true to type |
| Off-flavour | None | Strong (rancid, musty, scorched) |
| Moisture perception | Bone dry | Wet, weeping |
Score each sample in triplicate across sessions, average, and compare. A radar-chart of these attributes against your gold-standard batch is the single most useful quality picture a small producer can own.
Hedonic testing: do people like it?
- The 9-point scale: dislike extremely (1) β neither like nor dislike (5) β like extremely (9). Report the mean and the distribution.
- Panel = target market: 30+ consumers is the practical minimum for stable means; farmers-market customers are a fine start but they are self-selected fans β note the bias.
- Add purchase intent ("definitely/probably would buy") and one open question ("what would you change?").
- Watch the middle: a 5.8 mean with all scores at 5β7 is a safe, boring product; a 5.5 mean split between 2s and 9s is a love-it-or-hate-it product β different strategies follow.
Sensory shelf-life: the endpoint test
- Store dated samples from one batch under real conditions (your storage, your packaging).
- Test at intervals β 0, 1, 3, 6, 12 months β using the difference test against the fresh reference plus key attribute scores.
- Define the endpoint in advance β e.g. "off-flavour β₯ 3, or significant difference from reference" β before you start, or you will argue yourself into extra months.
- Set the best-before from the endpoint, with margin. This is the evidence behind the date on your label.
Printable score sheet
π Sensory score sheet
| Attribute | Score | Notes |
|---|---|---|
| Colour intensity (grey β vivid) | ||
| Uniformity (variable β identical) | ||
| Firmness (snaps β pliable) | ||
| Chewiness (short β long) | ||
| Sweetness (none β intense) | ||
| Acidity (none β sharp) | ||
| Aroma intensity (flat β intense) | ||
| Off-flavour (none β strong) | ||
| Overall quality (poor β excellent) |
Five ways panels lie
- Non-blind samples β the brand or process name on the plate biases every score. Code everything.
- Untrained "trained" panel β without anchors and practice, scores drift. Train with known references (e.g. deliberately scorched vs ideal mango).
- Wrong temperature β warm fruit reads sweeter; define and control it.
- Discussion before scoring β one confident voice moves the whole panel.
- Testing only fresh product β the interesting data is at 6 and 12 months; schedule the storage tests.