# ML & Datasets

> Status: Exploring

Exploring how raw data becomes a reliable machine-learning dataset, and what has to be true before training starts.

## The work before training

Machine learning gets discussed in terms of models, but how reliable a result is depends heavily on the data it was trained and evaluated on. So the part I am working through is everything before training: defining what the model should learn, deciding what data that requires, and getting it into a state worth trusting.

It is an interesting problem because it sits across software engineering, data engineering, domain understanding, and experiment design at the same time.

## From a question to a versioned dataset

1. **Define**: State the problem before deciding what data it needs.
2. **Collect**: Gather examples that are actually representative.
3. **Clean**: Missing values, duplicates, outliers, inconsistent records.
4. **Annotate**: Labels or ground truth, manual or assisted.
5. **Check**: Quality, label consistency, balance, coverage gaps.
6. **Version**: A snapshot with its preprocessing recorded.
7. **Split**: Train, validation, and test, without leakage.

## Open questions

- How should the problem definition shape the dataset before collection starts?
- How do you decide which examples are representative enough?
- When should annotation be manual, assisted, or automated, and how is its quality measured?
- How should imbalanced classes be handled, and when does augmentation distort the problem?
- When should a split be random, stratified, grouped, or chronological?
- How do you keep dataset versions and preprocessing reproducible?
- How are bias and coverage gaps found before training rather than after?

## Current scope

The current scope is dataset preparation, quality checks, versioning, splitting, and reproducibility before model training begins.

## Topics

- Datasets
- Quality control
- Reproducibility
