Dominic Feron

How to Read a FinTwit Statistic

A practical way to decide whether a small-sample bullish or bearish market statistic deserves attention without pretending it proves the future.

Every week my feed produces another market statistic. Eight down days have happened only six times. A breadth reading has appeared eleven times since 1990. A certain combination of price, volatility and moon phase is now 9 for 9.

Some are bullish. Some are bearish. A surprising number contradict one another before breakfast.

What are we supposed to do with them?

Consider a recent example. QQQ had closed below its open for eight consecutive sessions. The historical table looked like this:

Date VXN 1 month 3 months 6 months 1 year
Dec. 20, 201831.008.81%18.58%24.57%40.24%
Nov. 3, 201622.361.48%10.54%20.73%35.75%
Jul. 2, 201517.663.72%-3.40%4.20%1.27%
Apr. 12, 201119.394.86%2.28%0.82%20.28%
Jun. 28, 201029.372.08%9.67%21.69%25.43%
Dec. 29, 200843.692.39%6.85%26.55%60.23%
Median3.06%8.26%21.21%30.59%

Source: OddStats. Historical returns shown in the original table; the current occurrence had no forward result yet.

Six previous occurrences. All six were higher one year later. The median gain was 30.59%.

Should we buy immediately?

Not so fast.

The invisible experiment

The arithmetic in the table may be perfectly correct. The problem is everything we cannot see around it.

Was the eight-day rule specified before anyone looked at the returns? How many neighboring rules were tested? Perhaps seven days, nine days, close below open, close below yesterday’s close, SPY, QQQ, six months or one year.

Were the dates chosen after the result looked attractive? Were failed versions discarded? Are the six events independent, or are several expressions of one market regime?

We are shown the winning ticket, not the number of tickets sold.

This matters more than it first appears. If one pre-specified pattern lands in the most extreme 1% of outcomes, that is unusual.

If someone tests 1,000 variations and publishes the best one, an extreme result is part of the job description. Research on financial anomalies has wrestled with this problem for years. A conventional significance hurdle assumes one clean test. Repeated searching demands a much higher one.

The tweet rarely tells us the search history. It usually cannot tell us whether the attractive result survived parameter changes. We also do not know whether the event definition used information available at the time, or whether the sample was filtered by today’s index constituents.

Those omissions do not prove the pattern is false. They prevent us from knowing how much evidence the pattern contains.

As a forecast, the statistic is unusable in its published form.

That does not make it useless.

Build the ruler before seeing the object

You do not need to recreate the author’s research process. You can ask a narrower question: how strange would this result be if the dates had no special predictive power?

For a one-year market claim, the basic ruler can be prepared in advance. Take every eligible trading day in the history of the same asset and calculate the following 12-month return.

Then repeatedly draw random samples of 5, 6, 7 and so on. Record the distribution of the sample mean, median and positive hit rate for each sample size.

The result is a sampling-error map. It tells you what impressive-looking small samples the market can produce without a special signal.

A few broad regime versions make the ruler better. You might separate high and low volatility, bull and bear trends, or deep and shallow drawdowns. Keep the list short and fixed.

If you manufacture twenty regimes after seeing the six dates, you have merely moved the data mining from the tweet to your spreadsheet.

Now place the posted statistic inside the relevant distribution. An outcome near the middle is ordinary. The 90th percentile is mildly interesting. The 97th percentile deserves more attention. The 99.5th percentile is a strong anomaly candidate.

Notice the wording. A result at the 97th percentile does not mean there is a 97% probability that the signal is real.

It means that, under your stated null model, 97% of comparable random samples produced a lower statistic. The unknown number of trials behind the original post still prevents a calibrated probability of truth.

This is the useful middle ground. You cannot certify the signal. You can measure how hard ordinary sampling error has to work to explain it.

Base rates have a sense of humor

Win rates cause the most trouble. Six wins out of six looks immaculate. It may also be routine.

Suppose the market’s unconditional chance of a positive one-year return is 83%. Even with no signal, the probability of six positive years in six independent draws is about 33%.

The perfect score appears roughly once in every three six-event samples.

The same 6-for-6 result would be far more striking against a 50% base rate. The numerator has not changed. The information has.

This is why “100% bullish” is not an analysis. You need the unconditional hit rate, the distribution of returns and the sample size.

A high win rate can also hide poor economics if the few losing cases are much larger than the winners. Direction and payoff are different questions.

How high should the bar be?

There is no sacred cutoff. The following is a conservative attention rule, not a law of statistics:

Sample sizeSuggested treatment
(n<5)Usually an anecdote; interesting only if the result is extraordinary and has a plausible mechanism
(n=5–7)Look for an outcome around the outer 1% of the relevant null distribution
(n=8–10)The outer 5–1% can justify further attention
(n=11–15)The outer 5%, together with meaningful effect size, can clear the attention threshold
(n>15)Sampling error improves gradually; unknown selection bias remains

Percentile alone is not enough. A result should also differ by an economically meaningful amount from the normal market outcome.

It should not disappear when the single best event is removed. Its occurrences should represent separate episodes rather than six overlapping dates in one crisis. If the dates cluster in one obvious regime, compare them with that regime, not with the entire history.

These are cheap checks. They do not turn a tweet into a backtest laboratory. They stop the most ordinary noise from renting space in your head.

My compact rule is this: an (n=5–15) statistic earns attention when it falls in the outer 5% of a relevant historical null distribution and carries a meaningful effect. It must not be driven by one observation or become ordinary after a basic regime adjustment.

Below eight observations, I would usually demand the outer 1%.

“Earns attention” is doing important work here.

What the statistic can say

With this ruler, you can answer one useful question:

How unusual would this result be if it came from one pre-specified test?

You still cannot answer the grander one:

What is the probability that the claimed effect is real?

The invisible search process stands between the two. No confidence interval calculated from the published 5–15 observations can recover information that was never disclosed.

The honest verdict is conditional: nominally unusual under a defined benchmark, and worth treating as a hypothesis, but not calibrated evidence of a durable effect.

Sometimes that verdict will send you back to scrolling. Occasionally it will make you pause.

That is enough. A FinTwit statistic does not need to become a trading system to be useful. It only needs to survive a fair comparison with the surprising patterns that random markets manufacture every day.

The historical sample can earn your attention. It cannot demand your trust. The final observation will arrive without a confidence label, and the market will have the last word.