Menu
SIPS InsightsLive Behaviour

How to Tell When a Live EA Is Behaving Differently From Its History

A practical, evidence-based way to compare live trading against a strategy's own backtest — without overreacting to one bad week.

SIPSALGO·8 September 2026·8 min read
A shaded expected behavioural envelope with one live trajectory line deviating outside it.

Every systematic trader eventually asks the same question about a live strategy: is this still behaving the way its backtest said it would? Answering that well is harder than it sounds — it requires comparing like-for-like conditions, waiting for enough evidence, and resisting the urge to draw a firm conclusion from one bad week. SIPS’s own Historical vs Live Behaviour framework is a useful educational model for how to approach this comparison properly, whatever tools you use.

Compare like-for-like historical windows

The first principle is comparing genuinely comparable things. A strategy’s live return over the last 30 days should be measured against how that same strategy performed, historically, over similar 30-day windows — not against its all-time average, and not against an arbitrary benchmark. This kind of like-for-like comparison is what makes a divergence meaningful rather than an artefact of comparing two different kinds of measurement.

Three dimensions worth tracking separately

Rather than a single pass/fail verdict, it helps to track a strategy’s live behaviour across at least three distinct dimensions:

  • Return — how live results compare to historically similar periods.
  • Drawdown — how the live account’s drawdown compares to historical drawdown over similar windows.
  • Trade Frequency — how often the strategy is actually trading live, compared to its historical pattern.

A strategy can diverge on one of these while remaining entirely normal on the others — treating them separately gives a more precise picture than a single combined score would.

Evidence requirements: Building Sample

Before drawing any conclusion, there needs to be enough live evidence to draw one fairly. A useful discipline is requiring a minimum amount of live history, live trades, and comparable historical windows before classifying a strategy’s live behaviour at all — treating anything short of that threshold as still “Building Sample,” rather than forcing a premature classification. Reacting to two or three live trades as if they were a verdict is one of the more common — and most avoidable — mistakes in monitoring live systems.

A plain-English classification scale

Once there’s enough evidence, a percentile-based comparison against historically similar periods gives a natural, plain-English scale: results clearly better than most comparable historical periods (Outperforming), broadly typical (Normal), somewhat below typical (Watch), and clearly worse than most comparable periods (Deviation) — with an escalated Action Needed state reserved for divergence severe or persistent enough to warrant closer attention.

The wording matters here, and it’s worth being precise about it: a result described as “worse than most comparable historical periods” is a statement about relative ranking, not a claim about the size of the difference. “Better than 31% of similar periods” means 31% of comparable historical windows produced a lower result and 69% produced a higher one — it is not a statement that the current result is numerically 69% worse than something.

What one metric moving doesn’t automatically mean

A single metric drifting into Watch or Deviation territory is information worth understanding, not automatic proof of strategy failure. Short windows are noisy by nature, and some drift relative to history is normal even for a strategy performing exactly as expected. The framework is designed to flag genuine, sustained divergence worth investigating — not to punish ordinary short-term variation.

It’s equally worth restating a point from Why Trade Frequency Matters When Building a Portfolio: higher live trade frequency than the historical pattern is not automatically read as an improvement, any more than lower frequency is automatically read as a problem. Both are simply changes worth understanding in context.

These are monitoring indicators, not investment advice

It’s worth being explicit: a Watch or Deviation classification is a monitoring signal describing how live results compare to history — it is not investment advice, a recommendation to close or keep a position, or any kind of guarantee about what happens next. It’s an input to a trader’s own judgement, not a substitute for it.

Where this fits in the SIPS workflow

This is a description of SIPS’s real, currently implemented Historical vs Live Behaviour feature — the classifications, evidence thresholds, and Return/Drawdown/Trade Frequency structure described above match what’s genuinely shown once an MT5 account is connected. See Historical vs Live Behaviour for the full detail, Dashboard for where it first appears, and Live Portfolio Detail for the strategy-level view.

The practical takeaway

Comparing live behaviour to history well requires like-for-like windows, enough accumulated evidence before judging, and separate attention to return, drawdown and frequency rather than one blended verdict. Read carefully, this kind of comparison is genuinely useful early-warning information — read carelessly, it can just as easily produce false alarms or false reassurance.

Get started

Turn research into
better portfolio decisions.

Software and risk notice. SIPSALGO provides software tools for strategy and portfolio analysis. Trading and investment decisions involve risk, and analytical tools cannot guarantee future performance. Nothing on this page is financial advice or a recommendation to trade.