Executive summary
ToneBot is a systematic trading engine: a rules-based, long-only momentum and sector-rotation strategy, executed automatically under fixed risk controls. It invests in the parts of the stock market that are gaining strength and steps aside from the parts that are losing it, every trading day.
Markets trend. When an industry begins to lead — chip makers during an AI build-out, banks in an expanding economy, homebuilders when rates fall, gold in uncertain times — that leadership tends to persist for months, because large investors reposition slowly and new information takes time to be priced in. This tendency, called momentum, is among the most extensively documented effects in financial research.
ToneBot measures which areas lead, holds the strongest of them while they continue to lead, and rotates when leadership changes. Every rule is tested on years of market history it was not built on before it is used, and results are published, losses included.
Figures at a glance
| Official paper test | real broker, practice money | Oct 12 – Jan 18 |
|---|---|---|
| Paper account since the start | updated weekly | -0.1% |
| Strategies tested side by side | the official rules and challengers | 8 |
| Automated self-checks | run before every software change | 170 |
Practice and paper money only until a public go / no-go decision on January 18, 2027. Nothing in this paper is investment advice.
Methodology
The system follows the same five-step loop every trading day. It does not improvise and it does not act on tips.
Read the market
Each area of the market is scored on how strongly, and how steadily, it has outperformed the overall market. The system also reads conditions: whether the market is trending or choppy, whether interest rates are a tailwind or a headwind, and whether fear is rising.
Select only from leaders
New positions come only from leading areas, and only from investments already in their own uptrend with enough trading volume to enter and exit cleanly.
Size by risk
Position size depends on how much a position could lose, not on how attractive it appears. More volatile holdings receive smaller allocations; leveraged funds are treated as the larger risk they carry.
Exit by rule
Every position has an exit plan before it is bought: a protective stop that rises as the position gains, and a time limit to prove itself. Gains are allowed to run; losses are cut early.
Review
After the close the system grades its own decisions, compares actual fills with the plan, and records what it would change — without changing anything until a scheduled checkpoint.
Trading follows the left column every day. Learning runs continuously on the right, but it can only change the rules at a scheduled checkpoint.
A system versus a human trader
Neither is better at everything. The system is built to do what people do badly, and its known weaknesses are watched rather than hidden.
Where the system has the edge
- Discipline. Every rule is applied every time. No fear after a loss, no greed after a win, no revenge trades, no hesitation on a stop.
- Attention. It checks the market every ten minutes through the session, across more than fifty investments in thirteen themes, without tiring.
- Risk limits that can't be talked around. The daily circuit breaker, position caps and broker-held stops act automatically.
- Evidence before belief. Every rule is tested on years it was not built on, every trade is graded, and every forecast is scored.
- A full record. Each decision is logged with its reason, so mistakes can be found and fixed rather than remembered selectively.
Where a human has the edge
- Context. A person can read what the data doesn't show: the meaning of a policy shift, a rumour, a once-in-a-generation event.
- Adapting to the new. Rules learned from the past can fail when markets change character. The system adapts only at scheduled checkpoints, on purpose.
- Sharp reversals. Momentum strategies are known to suffer when leaders suddenly reverse. Size limits and the circuit breaker soften this; they don't remove it.
- Repeating a mistake. A flawed rule is followed faithfully until it is found and changed at a checkpoint.
- Dependence on machinery. Data feeds, brokers and servers can fail. Backups, failover and self-checks reduce this risk; nothing eliminates it.
Backtests can also flatter a system. Ours have been corrected twice so far: removing investments added to the list only after they had risen, and adding the real costs of leveraged funds to historical rebuilds.
Research and learning
ToneBot studies every week, but its trading rules change only at scheduled checkpoints. That separation keeps the public test honest while the research continues.
| Latest weekly experiment cycle | 2026-W41: 14 inconclusive | 4 kept, 14 rejected |
|---|---|---|
| Experiments run in the last 7 days | on years the rules were not built on | 14 |
| Research notes | plus new papers every Saturday | 6 |
| Trade forecasts graded | 0 awaiting their result after twenty trading days | 0 |
Latest week: 4 kept, 14 inconclusive, 14 rejected out of 32 ideas. Most ideas fail; that is the point of testing them.
Findings
- 2026-10-10
RiskHonest drawdowns are bigger than they first look
Our older backtests only counted a loss when a trade closed. Counting every open position at every day's close shows the worst drop of the current rules in some past crashes was larger than our own 35% limit. We now measure everything that way, and no change can be promoted unless it stays inside the limit across every historical period we test.
- 2026-10-10
RiskThe fix for big drops is sizing, not prediction
Across published research and our own tests, the changes that cut the worst drops were about HOW MUCH to hold in certain conditions, not about predicting the next move. One combination of sizing rules kept every tested period inside our limit and did better in most of them. It still has to pass our stress test before it is a candidate for January.
- 2026-10-10
ForecastingNobody (including us) can call one stock's next month
We tested 'look-alike' forecasts on 7,732 past cases across 12 well-known stocks and funds. Guessing whether a single stock would be up or down 20 days later was no better than a coin flip. The forecast RANGES, though, were well calibrated: about 79% of outcomes landed inside our 80% range. So the bot uses forecasts to size risk, never to pick winners.
- 2026-10-10
StrategyBolder is not better
We tested five aggressive or contrarian alternatives side by side (leveraged trend-following, all-in on the hottest sector, earnings-gap buying, buying fear, fading dips). None beat the current approach on recent years, and the most aggressive one nearly wiped out in the 2000-02 crash. The edge is in discipline, not in swinging harder.
- 2026-10-10
ReliabilityA trading bot must survive its own failures
We audited every way the system could go quiet or need a human: data sources, the broker login, a sleeping computer, alert channels. Each hole got a full fix — stale prices block new buys, alerts fall back to a phone push, two machines can never trade the same account at once, and a backup machine can take over.
- 2026-10-10
LearningIt learns every week, but changes only at checkpoints
During the test the trading rules are frozen so the result is honest. The bot still records what it WOULD change, grades its own forecasts and trades, runs experiments on years it has never seen and reads new research every Saturday. Only scheduled checkpoints can switch anything on.
Where attention went
Every time the system or its research desk reads, measures or decides something, it is logged. Activity over the last 7 days, by area:
| Individual securities | 254 | |
|---|---|---|
| Strategy | 185 | |
| News | 162 | |
| Founder input | 160 | |
| Risk | 110 | |
| Execution | 71 | |
| Market | 53 | |
| System health | 42 | |
| Quantitative models | 40 | |
| Experiments | 40 | |
| Research | 39 |
Share of the work: the automated system 273, the research desk 60, the founder 131.
Validation and performance
The pass marks were written before the test began, so they cannot be moved. On January 18, 2027 each is scored publicly. Real capital is committed only if all nine pass, and then at a quarter of normal size, scaling up over several months.
Since Oct 2: ToneBot −0.1%, S&P 500 +1.2% (behind by 1.3 points). Practice account, measured at each day's close.
Worst drop so far −3.0%. The test's pass mark is a worst drop smaller than 20% — far below this chart's range, measured every day, open positions included.
| Test | Pass mark | |
|---|---|---|
| 1 | Return versus the market | Beat the S&P 500 after all costs over the same dates. |
| 2 | Worst decline | Keep the largest fall from a peak under 20%, measured daily. |
| 3 | Execution quality | Real broker fills within a fraction of a percent of the planned price. |
| 4 | Model fidelity | The broker account behaves like its practice model. |
| 5 | Operational safety | No safety stop caused by a software fault and no unexplained order. |
| 6 | Reliability | Online for at least 95% of market hours. |
| 7 | Persistence of the edge | The advantage still appears on the newest data, not only in older years. |
| 8 | Readiness | Retirement-account mode built and rehearsed end to end. |
| 9 | Sufficient evidence | At least 30 completed trades. |
The weekly scoreboard begins once the test is under way.
A note on honest measurement
Our historical tests value every open position at every day's close, rather than recognising a loss only when a trade is closed. Measured this way, the current rules' worst historical decline exceeds our own limit in some past crises. Reducing it is the research desk's first priority, and no change is accepted unless it remains within the limit in every historical period tested.
Practice comparison since inception: strongest strategy +3.3%, weakest -2.2%.
Risk management
| Daily circuit breaker | A sufficiently bad day halts new purchases for the rest of that day. |
|---|---|
| Order controls | Every order is checked against the live price and a rate limit before it is sent. |
| Broker-held protection | Protective stop orders are held at the broker, so positions remain protected even if the system is offline. |
| Single active system | Each machine tags its orders; if two are ever active at once, one stands down automatically. |
| Self-monitoring | Stalled jobs, lost connections and stale prices are detected and repaired or escalated; stale prices block new purchases. |
| Retirement-account controls | Settled cash only, no short selling, a mandatory pause after a large decline until a person resumes trading, and a gradual scale-in. |
Safety settings are always active and are not altered during the test.
Research library
Topics studied by the research desk, newest first. Published research is a source of ideas; nothing is adopted without passing our own tests.
| 2026-10-10 | Research: Strategies with Verified Track Records |
| 2026-10-10 | Self-reliance holes — audit Oct 10, 2026 4:50 PM |
| 2026-10-10 | Moon Shot lanes — first scorecard |
Recent papers under review
- Continuous Timing Signals for Growth-Defensive Style Allocation: Factor Attribution, Risk Matching, and Out-of-Sample Evidence
- Forecast-to-Fill: Benchmark-Neutral Alpha and Billion-Dollar Capacity in Gold Futures (2015-2025)
- End-to-End Parametric Portfolio Policies for Cross-Asset Futures Timing: When Do AI Models Beat Simple Rules?
- Observable Matrix Dynamics of Stocks
- (Non-Parametric) Bootstrap Robust Optimization for Portfolios and Trading Strategies
- Deep Hedging with Reinforcement Learning: A Practical Framework for Option Risk Management
- Regime-Conditional Distributional Comparison of Trading Strategies: A GAMLSS/ZAGA Framework Applied to the S&P 500
- Refining and Robust Backtesting of A Century of Profitable Industry Trends
Build effort and cost
What this system would take to build by hand, set against what it has actually cost.
The actual spend is about 0.2% of the low estimate — too small to see at this scale.
| Software | 45,039 lines of working code across 119 modules |
|---|---|
| Automatic checks | 170 self-tests, all run before every release |
| Releases | 111 logged improvements since September 30, 2026 |
| Founder | At least 56 hours over 12 days, measured from activity logs. 76 standing rules set, 147 ideas submitted, 105 of them built. |
| Paid | Claude Max plan (AI assistant) $100/month since September 2026; Cloud server $5/month since October 2026 |
| Free | market data (Yahoo Finance + Schwab), paper broker (Alpaca), website (Cloudflare Pages), alerts (Discord), research feed (arXiv) |
The system has cost roughly 0.2% of what a hand-built equivalent would. It was designed and directed by its founder — every rule, idea and approval above — and written with an AI assistant (Claude), which is where nearly all of the difference comes from. Founder time is a floor: most time spent in planning conversations leaves no trace in the logs.
How the estimate is made
Only lines of working code are counted (no blank lines or comments). We assume an experienced developer finishes 150–300 tested lines per 8-hour working day and charges $100–150 per hour, common figures for contract financial-software work. Research, design and the time spent testing strategies are left out, so the true hand-built figure would be higher. The count updates with every edition.
Revision history
Improvements to the system, newest first. Strategy changes occur only at checkpoints; most entries concern measurement, safety and reliability.
| 2026-10-10 | Professional white paper + faster terminal + MIND button |
| 2026-10-10 | Public white paper in #how-it-works |
| 2026-10-10 | Mind crawler — watch the bot and Claude think |
| 2026-10-10 | Honest numbers everywhere |
| 2026-10-10 | Predict, then grade: per-trade forecasts + crash-risk lab |
| 2026-10-10 | Real-money groundwork + cleanup decisions |
| 2026-10-10 | Scout: research inbox + pros' momentum funds (no AI, read-only) |
| 2026-10-10 | Terminal: calmer vitals page |
| 2026-10-10 | Self-reliance round 2: two-machine safety, failover, better Friday recap |
| 2026-10-10 | Self-reliance fixes: fewer ways the bot can go quiet or need Claude |
| 2026-10-10 | Weekly forecast + calibration (…) |
| 2026-10-10 | Trade-quality grade (…) + roadmap dates |
| 2026-10-10 | Dead holdings + quieter logs (…) · Self-tuning race account (…) |
| 2026-10-10 | Self-tuning frozen for the paper test (…) |
| 2026-10-07 | Official paper test moves up to Fri Oct 9 |
| 2026-10-07 | Risk desk: how many REAL bets each account holds |
| 2026-10-07 | HEALTH light in the terminal + morning pre-flight |
| 2026-10-07 | One health check for the whole bot |
| 2026-10-07 | Secret: the all-seeing eye (idea #140) |
| 2026-10-06 | SSD backup follows a renamed drive |
| 2026-10-06 | Terminal: background code removed |
| 2026-10-06 | Vertical terminal: roomy layout, full crash meter, bigger trade journal |
| 2026-10-06 | Discord clean-up: #how-it-works is one card, paper button fixed |
| 2026-10-06 | 11:45 check: idea box quieter, #general duplicates stopped, phone terminal fixed |
| 2026-10-06 | Safety net at the broker + investor kit for Wednesday |
| 2026-10-05 | Terminal polish for Wednesday + Strategos branding + investor-friendly Discord |
| 2026-10-02 | Safety pass before paper trading + a big strategy lab |
| 2026-10-02 | ToneBot Terminal app · Monday rehearsal passed |
| 2026-10-02 | ⏪ Replay · briefing · compare chart · wall mode |
| 2026-10-02 | Paper graphs + tap-for-positions |
| 2026-10-02 | Email: weekly only |
| 2026-10-02 | Race scoreboard |
| 2026-10-02 | Trade journal, new channel layout, expected $ |
| 2026-10-02 | Discord tune-up: less noise, more fun |
| 2026-10-02 | Paper trading wired into everything |
| 2026-10-02 | Terminal fits any screen · boot-up + command line |
| 2026-10-02 | New weekly recap · terminal race board |
| 2026-10-02 | Legends lab + channel check |
| 2026-10-02 | Pick review: why + how we'll improve · where the math comes from · savings |
| 2026-10-02 | Pick review + weekly summary |
Glossary and disclosures
| Systematic trading | Trading in which fixed, pre-tested rules make every decision to buy, sell or size a position, rather than a person deciding case by case (discretionary trading). ToneBot is systematic, quantitative (its rules are built and tested on data) and algorithmic (software executes the orders). |
|---|---|
| Momentum | The tendency of investments that have been rising to keep rising for a period. |
| Sector rotation | Moving capital between areas of the economy as leadership changes. |
| Drawdown | The decline of an account from its highest value. |
| Backtest | Applying the rules to historical prices. |
| Walk-forward test | Choosing settings on older years, then evaluating them on newer years they were not built on. |
| Paper trading | A real broker account with real prices, funded with practice money. |
| Calibration | When the system says an outcome is 80% likely, it should occur about 80% of the time. Forecasts are graded on this. |
Disclosures
This paper describes a system under test. It is not an offer, a solicitation or investment advice, and it should not be used to make investment decisions. Historical and simulated results have inherent limitations and do not guarantee future results. Live results may differ materially. The system's specific rules, parameters, holdings and investment universe are proprietary and are intentionally withheld.