← All work

Case study

CalmTrader

An active agentic investment-research project building auditable workflows for policy, evidence, market structure, and evaluation.

Role
Co-Founder / CTO
Period
May–Dec 2025; Jul 2026–present
Status
Active private R&D

CalmTrader is a research project. Nothing here is investment advice.

What we are building now

CalmTrader is active again. We are building an agentic research and decision-support system that turns an investment objective into an auditable chain of policy, evidence, constraints, and judgment.

The goal is not to hand a model a portfolio and accept whatever it recommends. We want each consequential step to be inspectable: where the data came from, which constraints were applied, what the research agents found, which claims they could support, how model judgment was evaluated, and where deterministic code has the final word.

My role now

CalmTrader is a working group rather than an incorporated company. I build it with Navaratnasothie (Selva) Selvakkumaran, Ph.D.. I am Co-Founder and CTO, and I returned in July 2026 after a six-month break for full-time parenting.

While I was away, Selva built the first working implementation of this new direction. I am now leading the technical direction with him: turning that foundation into an environment we can use continuously, deciding which agentic patterns belong in the product, and connecting the research loop to measurable evidence.

The next milestone

Our next milestone is to turn the current prototype into the working research environment we use day to day. That means connecting live inputs with persistent context, making the policy and research stages repeatable, evaluating both deterministic and model-based behavior, and using what we learn to decide which capabilities should become a public product.

We are deliberately keeping brokerage integration and order placement outside this product generation. The near-term objective is a trustworthy research system with a clear evidence trail, not a public advisory or managed-trading service.

The foundation we are building from

Selva’s prototype turns a plain-English investment goal into a schema-validated policy, generates candidate ideas with citations, enforces hard portfolio constraints in deterministic code, and asks a model for qualitative scoring only after those checks pass. A parallel group of research agents produces cited memos, identifies unsupported claims, and exposes its tool calls and results for inspection.

It also includes gamma-positioning analysis with explicit freshness and data-provenance indicators, plus an evaluation harness for policy parsing, tool contracts, market analytics, and memo quality. The distinction between deterministic controls and model judgment is intentional: a model can help interpret evidence, but it cannot override a numerical risk rule.

Selva built this iteration through an autonomous Claude Code workflow organized around a written backlog, persistent state and decision records, and a verifier agent with binding acceptance checks. This iteration was his work; it is the starting point for what we are now building together.

Milestone one: research through execution

The first CalmTrader system connected strategy development, backtesting, risk controls, simulation, and live futures execution. I led that technical work, wrote most of the code, and performed much of the backtesting. Hayk G. was a significant engineering contributor and worked with me on implementation and operations.

We used TradingView and Pine Script to turn trading ideas into explicit strategies and test them against historical data. The work covered EMA market structure, RSI and MCDX reversals, VWAP behavior, ATR and volatility, volume, price zones, fair-value gaps, and session and liquidity structure.

Once a strategy was enabled, the execution system consumed its signals, reconciled them with account and position state, and routed eligible orders to Tradovate. It could run against either a simulation broker or a live account. Risk controls included separate demo and live settings, profit and loss limits, account pauses, broker-enforced liquidate-only states, stop and exit handling, duplicate-signal protection, and a record of trades that were executed or prevented.

That milestone proved that research, strategy definition, risk policy, simulation, execution, and reconciliation could share one architecture without becoming one opaque process.

Why it remained private

Our original ambition was a public SaaS product. We narrowed it to private research and personal trading once we recognized that personalized futures guidance or automated trading for other people could cross into regulated advisory activity.

We could not responsibly handle that with a disclaimer or as a feature added near launch. Determining where the product belonged would require legal analysis, registration or exemption decisions, operating controls, and continuing compliance work. That was too large a burden for a small team, so we kept the system private and used the experience to define a safer product boundary for the next milestone.

How we are doing the research

Loop engineering makes the full idea-to-backtest cycle explicit: propose a hypothesis, encode it, run it, examine the evidence, and use the result to choose the next experiment. A graph of specialized agents can strengthen that loop by screening an idea before it consumes a full coding-and-backtest cycle: challenging its causal logic, checking whether something similar has already been tried, estimating whether costs would erase the edge, and judging whether the specification is concrete enough to implement.

Deciding which agentic patterns belong in CalmTrader means building them somewhere I can measure them first. I validated the critic-graph half of this idea in a separate trading-research sandbox of my own, before proposing it for the product. Each strategy idea there passes through four critic agents: causal mechanism, redundancy, cost realism, and feasibility. An orchestrator samples the two or three that matter for a given idea, weights each critic by its own confidence, and resolves disagreements with a single bounded round of message passing instead of an open-ended debate. The pattern is adapted from a 2026 Graph-of-Agents framework.

The part I care about most is how a verdict earns trust. The first live round ran in shadow mode: every triage call was computed and logged, but nothing was pruned, until those calls could be compared against real backtest results. Only a verdict that survives that check is allowed to filter ideas for real. That sandbox is not CalmTrader’s production pipeline, and I keep the two separate on purpose. But it is how I answer, with evidence instead of assertion, whether a pattern like this deserves a place in CalmTrader’s research loop.

What we are testing now

  • Whether parallel research agents can produce conclusions that remain sourced and inspectable
  • Where deterministic constraints should overrule or bound model judgment
  • How evaluation and visible provenance change the trustworthiness of an AI research product
  • Whether the critic-graph triage I validated in a sandbox improves experiment quality inside CalmTrader’s own loop, not just in isolation
  • Which lessons from private research can become a useful, responsibly bounded public product