Im just pasting a post I made on another thread in the reviews section (Red 707) because it gives more context to what we are discussing on this thread. See below:
707 was the same for everyone, and we pretty much got all enough survey data to make a decision before anyone even started talking about it on here. It's officially going to make the menu at some time in the future, but we havent even began to consider when that might be. However, I have been getting a lot of inquiries about the complexity of a testing process, so Ill give a sneak peek of 707 since we are on the subject

Note that this is just part of the survey data, I'm not giving the full picture here to try and keep it breif but I think its enough to illustrate what I'm about to break down.
Just looking at the survey data of this batch, you can see that its average overall quality score is 8.22 (on a scale of 1-10). When we first began our research, the average score was basically all we looked at. I mean it makes sense, right? However, just looking at the average score is actually a lazy and ineffective way to use this data to evaluate overall quality - there is a lot more that must be considered. There are even other steps that come between collection and analysis that should be considered such as “survey cleaning”, but that’s a deeper rabbit hole that I wont go down here for the sake of brevity (who am I kidding in thinking this might be brief, lol). When reading survey data, if the details are ignored or the data is read poorly, it can really limit our ability to capture valuable insight and it dramatically reduces the credibility of our findings. Over the years we have become much more refined in how we collect and interpret data, I’ll give just a few high-level examples of some other ways to view the data.
Now back to the overall quality score of 8.22. On the surface that doesn't sound very impressive. It’s not bad, but its not great either. However, you have to first consider what the scale of 1-10 represents.
1 = Poor Quality
5 = Average – comparable to what you will find from most online vendors
10 = Exceptional - some of the best leaf currently available.
Following that scale, if a "5" is average, then "8.22" is more than above average, it's actually creeping into the range of really good. However, that's just part of the story that is being told here.
Our experience has taught us that you will hardly ever find a batch that is a home run for every person. Thanks to our "friend" personal chemistry, its exceedingly rare that we find a batch where 100% or even 95% of testers score it above a 7 or 8.... it just doesn't happen very often. Since our goal is to construct a menu where the consumer will on have a higher probability of finding batches that work well for them, it's really important for us to take into consideration the overall approval ratio of each batch.
A score of “7” is generally defined as “above average” and most consumers give “menu approval” to items they score 7 or higher. Out of 32 surveys, only 3 people scored it below a "7" with the lowest score being a "5" (6, 6, 5). Accepting that standard, this batch has an approval ratio of 29/32, or roughly a 91% favorability rate.
I know that 91% doesn’t sound like anything special if this were ratings on Amazon…it would basically equate to a 4.5 of 5 stars. Whatever the product was, that would be good enough for me to feel comfortable pulling the trigger, however, it wouldn’t give me the impression that this thing is particularly awesome. However, in the world of kratom (or “kratom roulette”), having 9 out of 10 people consider the batch as “above average” it’s actually pretty good. Back in my days as a consumer, if I were told I have a 90% chance of being pleased with a batch it would absolutely make my list of batches to order.
However, the story is still not complete. Let’s take a look at how the scores of those who approved are broken down:
10 = 7 respondents
9 = 7 respondents
8 = 8 respondents
7 = 7 respondents
We can see here that of those who liked this batch, roughly 50% of surveyors found it to be one of the best batches available or very close to it (9 or 10). The other 50% found it to be above average/good (7 or 8). So if we break down the probabilities that a person will like it, we have the following:
0% will find it as poor quality
9% will find it as average
46% will find it as good/ very good
44% will find it as the best you can get or very close to it
Those odds are pretty promising when you break it down like that. As a consumer, when its broken down this way it is much more appealing than if I saw it had an average score of 8.2. We are in the business of providing what the consumer wants, so its really important that we look at things this way.
But that’s not even where the analysis ends. There are other ways to analyze the data such as seeking out a central tendency by removing the highest and lowest 10% of survey scores (this can help offset some of the one-offs that might exist because of a tester having an "off" day). There are tons of different techniques to apply here, and they can all help give you a better picture.
Also, remember that I am only sharing part of the data here. There are other things I can consider here such as providing different “weights” to each individual surveyor's score. Perhaps I want to decrease the weight of surveys from people who have been consuming kratom for less than one year, and give more weight to surveyors with 5+ years of experience. I can also assign weight by a surveyor's average score or even cross-reference it with how a surveyor has scored other high-quality batches in the past. I could also look at the surveyor's other history to remove “straight line” answers. I could also begin breaking down correlations with the preference profiles of individual surveyors. I'll stop with the possibilities, but as you can imagine, sometimes it pays to get into the nitty-gritty especially if we are having a hard time making a decision.
Another thing to consider is the possibility of just having “bad data”. I always tell surveyors that I would rather have no data than bad data, but it still happens. For example:
- There are people who might have submitted 7 surveys at once and they are just assigning scores to each sample by memory (this usually does not provide good data).
- Sometimes scoring from one tester may appear to be random or contrary (such as straight-line answers, or answers that with each sample contradict the crowd). This could indicate that the surveyor was just submitting them to check it off as complete, or perhaps they are fighting a virus or something and just can't get a good read on things.
These are just a couple of things that should be taken into consideration, especially if you don’t have a very big pool of data to play with. Occasionally we might see a great divide between the people who absolutely love a batch and those who score it a 1-3. I'm looking at a real-life example right now where we only have 20 surveys in and 15 people (75%) rated it between “7-10” and 5 people (25%) rated it a “3” or below. The average turns out to be 6.65, which right away looks like it’s not menu worthy. However, with a good chunk of people loving it and such a small pool of data, it's hard not to wonder if maybe 1 or 2 of those who scored it low were just not primed for kratom that day and no matter what they tested, it was going to be a “1”. If we were to just turn 2 of the low scores into an “8”, well suddenly the overall quality score is a “7.25”… that getting in the territory that might need to expand testing or look at the data in a different way (especially since some people really seemed to like it).
In the scenario above, we might decide to simply expand the testing pool or maybe target those testers who scored low with the same sample in the future, but under a different tester ID. A lot of you have noticed that sometimes it takes a long time for a batch to appear on the menu, this can sometimes be the reason (there are other reasons too that have to do with a planned release schedule to ensure our batches are consistent year-round).
Im finding myself getting into territory where I could talk forever here, so I am just going to wrap up. It wasn’t my intention to get this deep, but I think there are likely one or two fellow nerds here who might appreciate it.
