Sample A or Sample B? How to Run a Fair Comparison

By admin
The short answer

Choosing between two samples is a decision, not a preference, and it deserves the same treatment as any other decision: the same conditions for both options, criteria agreed in advance, and a written record of the outcome. Most brands compare samples in a room, at different times, on different days, and then wonder why the team cannot agree. Set the protocol first and the argument usually disappears.

Sample A or Sample B? How to Run a Fair Comparison——全文要点速览

Key takeaways

  1. Compare like with like: the same concentration, the same fill format and the same testing conditions, or the result measures the wrong variable.
  2. Agree the criteria before you smell anything, so the comparison is not rationalised after the fact.
  3. Blind the samples and repeat the test at a second session; a preference that survives a repeat is a finding.
  4. Score the characters that must be present and the ones that must not be, rather than scoring overall likeability.
  5. Write the decision and the reason down, because the next revision depends on why the losing sample lost.

Two samples, one decision, and a team that cannot agree. This is a normal situation in fragrance development, and it is almost never a problem of taste. It is a problem of method.

When one sample is sprayed at nine in the morning on paper and the other at four in the afternoon on skin, the comparison already has a built-in answer that has nothing to do with the fragrances. The same is true when one sample has rested for a month and the other was filled yesterday.

The method below is short. It takes two sessions and a scoring sheet, and it removes most of the noise from the decision.

First, make the comparison valid

Validity comes before judgement. Both samples must be at the same concentration, in the same format, applied the same way, and read at the same intervals. If one is a trial built at a higher dosage than the other, the comparison tells you which dosage you prefer, not which direction.

It is worth knowing how concentration changes what a fragrance does before scoring anything. how fragrance concentration changes performance explains why the same direction can read rich at one level and thin at another, which is precisely the variable that wrecks an A-versus-B session.

If the comparison is between suppliers rather than between two formulas, keep the method and change the criteria. Ask each supplier the same written question and score the answers the same way; a service page such as the one published by this Chinese perfume manufacturer gives you a fixed reference for what a complete answer looks like.

Same day, same hour, same nose

Test both samples in the same session, on the same person or the same panel, with a break between them. The nose adapts quickly, so the second sample should not be judged by a nose that has just spent ten minutes on the first one.

If the panel has to split across days, split it in a way that both samples are tested by the same people under the same conditions, and compare the readings within each tester rather than across testers.

Blind the codes and keep them blind

Label the vials with neutral codes and keep the key with someone who is not doing the smelling. A sample that is known to be the new revision has an advantage that has nothing to do with its smell.

A scoring sheet that produces a decision

CriterionSample ASample BWhat the score means
Opening character1-51-5Whether the first fifteen minutes match the brief
Development at 30 minutes1-51-5Whether the middle holds the intended direction
Base at 3 hours1-51-5Whether the dry-down matches the price positioning
Projection and sillage1-51-5Whether it fills the intended use occasion
Longevity1-51-5Whether it survives a working day or an evening
Fit with the brief1-51-5Whether the exclusions were respected
Must-not-be-present charactersPass or failPass or failA hard gate, not a score

Keep the last row separate from the scores. A sample that hits every note of the brief but carries a character the brand explicitly ruled out is not a close second; it is the wrong sample, and averaging the scores would hide that.

Illustration: A scoring sheet that produces a Decorative illustration for the section "A scoring sheet that produces a"; visual only, carries no data.

What to do with a split panel

If half the room prefers A and half prefers B, that is useful information rather than a deadlock. Ask the two groups to describe what they smell differently; a fragrance that reads clean to one person and sharp to another is telling you something about the direction, and it will behave that way in the market as well.

A split panel is also a signal to look for the third option. When two directions each satisfy part of the brief, the answer is often a revision that takes the successful element from each, which is cheaper than two more rounds on either.

Score the constraints before the preferences

Technical limits come first. A material's permitted use level is a separate question from how it performs in a blend; the industry's own framing of safe use is a scientific assessment rather than a sensory one [1]. Score the compliance pass or fail, then the preference.

Do not average away a disqualification

A sample can win on six criteria and still be unusable because it fails a seventh that the brand cannot compromise on. Decide which criteria are gates and which are scores before the session begins; the team will otherwise do it unconsciously, with the loudest voice deciding the weighting.

The protocol, in order

  1. Confirm both samples are at the same concentration and were filled at comparable dates.
  2. Agree the criteria and mark which ones are pass-or-fail gates.
  3. Code the vials and keep the key out of the room.
  4. Apply both samples to paper and to skin, and write the time on each strip.
  5. Read both at the opening, at thirty minutes and at three hours, recording each reading separately.
  6. Repeat the whole session on a second day with the same panel.
  7. Compare the two sessions, then write the decision and the reason for it in one paragraph.
Illustration: The protocol Decorative illustration for the section "The protocol"; visual only, carries no data.

A comparison that cannot be repeated is an opinion. A comparison that produces the same answer twice, on coded samples, with the criteria recorded before the session, is a decision the whole team can defend — including to the perfumer who has to act on it.

Turn the result into a brief for the next round

The losing sample is not wasted work; it is a map of what to avoid. Write the next brief as a list of things to keep from the winner and things to remove from the loser, and name the criteria where neither sample was good enough.

Illustration: Turn the result into a brief for Decorative illustration for the section "Turn the result into a brief for"; visual only, carries no data.

If the comparison is being run to choose between suppliers rather than between formulas, the same discipline applies, and it is worth seeing examples of fragrance projects to understand what a finished brief looks like before you write the next one. A supplier who can work from a structured brief and return a structured answer is easier to manage than one who needs to be guided through every round.

One last habit: keep both samples after the decision. The rejected direction often becomes the starting point for a second product a year later, and a retained sample is the only reliable way to return to it.

Sources

  1. IFRA: Safe Use and Fragrance Science —— IFRA's explanation of how fragrance materials are scientifically assessed for safe use and how those conclusions are applied by the industry.

Frequently asked questions

How many samples can a panel evaluate in one session?

Two or three. Beyond that, adaptation and fatigue start to distort the readings, and the fourth sample is effectively being judged by a tired nose. Run repeated sessions rather than long ones.

Should samples be compared on paper or on skin?

Both, for different reasons. Paper is consistent and can be kept for reference; skin shows how the fragrance performs with warmth and time. If the panel compares on skin, make sure both samples go on skin, at the same sites.

What if the two samples were made at different times?

Note it, and treat the difference as a variable. A sample that has rested longer may read differently from a fresh one, so either wait until both have settled comparably or accept that part of the difference is time rather than composition.

Is a split panel a bad sign?

Not necessarily. Different people perceive materials differently, and a direction that divides a small panel may simply be a stronger, more distinctive scent. Use the split to find out what each group is reacting to before deciding it is a problem.

How should the decision be recorded?

One paragraph naming the chosen sample, the criteria it won on, the criteria where neither sample was adequate, and the specific changes requested for the next round. That paragraph is the brief for the revision.