Lab experiment

ML & Datasets

Exploring how raw data becomes a reliable machine-learning dataset, and what has to be true before training starts.

Status
Exploring
Concepts
  • Datasets
  • Quality control
  • Reproducibility

This diagram is a working model, not a deployed architecture.

A dataset preparation pipeline under exploration: a problem definition sets data requirements, records are collected, cleaned of duplicates and outliers, annotated, and passed through a quality gate into a versioned dataset snapshot, which then separates into distinct train, validation and test partitions.
01

observations

The work before training

Machine learning gets discussed in terms of models, but how reliable a result is depends heavily on the data it was trained and evaluated on. So the part I am working through is everything before training: defining what the model should learn, deciding what data that requires, and getting it into a state worth trusting.

It is an interesting problem because it sits across software engineering, data engineering, domain understanding, and experiment design at the same time.

02

workflow

From a question to a versioned dataset

  1. 01Define

    State the problem before deciding what data it needs.

  2. 02Collect

    Gather examples that are actually representative.

  3. 03Clean

    Missing values, duplicates, outliers, inconsistent records.

  4. 04Annotate

    Labels or ground truth, manual or assisted.

  5. 05Check

    Quality, label consistency, balance, coverage gaps.

  6. 06Version

    A snapshot with its preprocessing recorded.

  7. 07Split

    Train, validation, and test, without leakage.

03

next questions

Open questions

  • How should the problem definition shape the dataset before collection starts?
  • How do you decide which examples are representative enough?
  • When should annotation be manual, assisted, or automated, and how is its quality measured?
  • How should imbalanced classes be handled, and when does augmentation distort the problem?
  • When should a split be random, stratified, grouped, or chronological?
  • How do you keep dataset versions and preprocessing reproducible?
  • How are bias and coverage gaps found before training rather than after?
04

observations

Current scope

The current scope is dataset preparation, quality checks, versioning, splitting, and reproducibility before model training begins.

Previous experiment
TASKATTEMPTVERIFYDONELESSONSPASSIMPROVE
Self-Improving AI Agents

Still exploring

Thinking about the same problem?

I'm always interested in what works, what fails, and how to test the difference.