← All work

Case study

ObexMetrics

An early machine-operable platform for persistent, parallel data-science workflows, now informing agentic and forward-deployed systems.

Role
Co-Founder / CTO
Period
Nov 2023–Apr 2025
Status
Completed product / active reference

Selected media

See the product in motion.

Platform overview Revolutionizing machine learning with the ObexMetrics platform An introduction to a more repeatable way to run and build on machine-learning experiments. Watch on YouTube
Product demo Gen-AI powered exploratory data analysis A demonstration of the exploratory-data-analysis assistant I built for data scientists. Watch on YouTube
Quick demo Exploratory data analysis with GPT-4o A shorter look at the ObexMetrics exploratory-data-analysis chat tool. Watch on YouTube

The problem

Data scientists often do their best work in notebooks, but notebooks are difficult to turn into shared institutional memory. After enough experiments, it becomes hard to reconstruct which data, transformations, algorithm, parameters, and results belonged together. Starting a new investigation can mean manually rediscovering work the team has already done.

We built ObexMetrics to give that work a durable home. The platform preserved the artifacts and lineage of each run, made earlier experiments available to teammates, and allowed many experiments to run in parallel rather than tying the process to one notebook session.

My role

My co-founder, Navaratnasothie (Selva) Selvakkumaran, Ph.D., architected the full stack and the MLOps core, set the technical vision and product strategy, and led customer discovery and fundraising. As Co-Founder and CTO, I worked hands-on across the interface and infrastructure and implemented the platform with the development team, including Hayk G.

My largest direct implementation was the EDA chatbot: a multi-user assistant that could work with a customer’s data, maintain a conversation, produce analysis and suggestions, and collect feedback on its responses.

What we built

The working product included:

  • Persisted datasets, configurations, experiments, runs, and results
  • Lineage and checkpoints that made earlier work recoverable
  • Collaboration and reuse across users
  • Parallel experiment execution and forecasting workflows
  • A forecasting assistant and an LLM-based exploratory-data-analysis assistant
  • Workloads that ran inside a customer’s own AWS account

For customer-cloud execution, the customer granted ObexMetrics an AWS role. The platform assumed that role through AWS STS, received temporary credentials, and created ECS/Fargate tasks in the customer’s environment. Compute-heavy work and its cloud cost stayed in the customer’s account, and ObexMetrics did not need permanent access keys.

ObexMetrics evolved through more than one product shape. An earlier version had working, deployable modules for tabular machine learning and computer vision. In February 2024, we deliberately retired those surfaces and narrowed the product around Algorithm Builder and Forecasting, including the EDA and forecasting assistants.

The other experiments did not reach that same level of maturity. A general-purpose LLM module was integrated briefly, while presentation tooling and a later fine-tuning module remained prototypes. That distinction matters: this was more than a collection of roadmap ideas, but it was never one uniformly released suite either.

Why it was early

ObexMetrics asked data scientists to move work out of familiar notebooks and into a structured platform before the agentic payoff was obvious. At the time, durable artifacts, lineage, explicit configurations, resumable stages, and shared execution could feel like extra process. Reproducibility, collaboration, and parallelism were not always enough to justify changing a workflow people already understood.

That calculation looks different once an LLM can operate the workflow. The structure was not incidental overhead; it was the state and tool surface an agent needs. A model cannot reliably continue a long-running investigation if the relevant dataset, transformations, parameters, results, and prior decisions exist only in a notebook and someone’s memory.

We were early in trying to make that work machine-operable. We did not yet have today’s orchestration patterns or dependable long-running coding agents, but we had already turned much of the data-science lifecycle into durable operations that an agent could eventually inspect and invoke.

An agent-operable platform

ObexMetrics was more than an experiment tracker. By turning datasets, configurations, algorithms, runs, results, lineage, and cloud execution into persistent, addressable state, the platform created a harness an AI agent could operate.

An agent could, with the appropriate tool layer, select a governed data artifact, configure an experiment, choose or build an algorithm, launch parallel runs, compare the results, and continue from a checkpoint without losing context. The released EDA and forecasting assistants demonstrated parts of that direction. Fully automatic discovery was not an end-to-end feature we shipped; it is what the underlying platform now makes plausible.

The EDA assistant

The assistant used AWS Lambda, the OpenAI Assistants API, TypeScript and Python services, PostgreSQL and DynamoDB state, and WebSockets. Sessions were tied to individual users and forecast files, and the workflow included response ratings and end-to-end tests.

The important step was connecting the conversation to a governed data artifact. This was more than placing a chat window beside a CSV: the assistant’s analysis became part of the same persistent experimental record as the rest of the work.

I documented the architecture in a three-part technical series beginning with Building a robust multi-user chat assistant.

What happened

Being early did not make the business successful. We did not find product-market fit or secure the funding needed to continue. During customer conversations and demonstrations, many practitioners preferred their existing notebooks and local files. Some also viewed automation as a threat rather than a useful extension of their work.

The adoption problem was partly one of timing: we were asking people to accept the cost of a structured workflow before agents made the benefit concrete. The underlying technical problem has not gone away. Long-running AI work still needs persistent context, checkpoints, parallel branches, evaluation, and a reliable way to resume after asynchronous work finishes.

What I am carrying forward

I am now revisiting the ObexMetrics thesis through independent work on agentic systems and my forward-deployed engineering capabilities. These projects are laboratories for learning how to enter a complicated domain workflow, make its state and decisions explicit, put useful tools under agent control, evaluate the result, and turn what works into reusable infrastructure.

That is also the forward-deployed pattern I see in ObexMetrics itself: work closely enough to understand the customer’s real environment, run securely inside customer-controlled cloud infrastructure, solve the immediate workflow, and feed what the field teaches back into the platform.

ObexMetrics is no longer an operating company, but its architecture remains an active reference for the systems I am building now.

What the project shows

  • Notebook experiments can become durable, reusable product state.
  • A serverless control plane can coordinate parallel ML work in customer-owned infrastructure.
  • An AI assistant becomes more useful when its conversations and outputs are connected to governed data.
  • Structured tools and persistent state can turn a human-operated workflow into one an agent can operate.
  • A sound technical thesis can arrive before customers are ready to change their workflow.