Volunteering / technical training and review
Volunteer roleAI for Impact
Technical Trainer, Deliverables Reviewer & Organizer
Role
- Title
- Technical Trainer, Deliverables Reviewer & Organizer
- Type
- Volunteer
- Training
- Preparing and delivering technical training to participants
- Support
- Technical support to teams during the event
- Review
- Analysis and review of technical deliverables
- Organisation
- Hackathon organisation
What is reviewed for
- Reproducibility
- Can someone else run this and get the same thing?
- Data methodology
- Where the data came from, what was done to it, and whether the split is honest
- ML methodology
- Whether the result is a result or an artefact of the evaluation
- Clarity
- Whether the deliverable states what it did and what it did not
01
The role
- Technical trainingPreparing and delivering training on the tooling and the method: how to set a project up so that it can be handed to someone else, how to handle data, and how to evaluate a model without fooling yourself.
- Participant supportTechnical support during the event itself, where the questions are concrete and the clock is running: an environment that will not build, a pipeline that leaks, a metric that looks too good.
- Deliverable reviewReading what teams hand in and assessing it technically, which mostly means asking whether the number in the summary survives someone else rerunning the notebook.
- OrganisationHackathon organisation.
02
What review actually catches
The failure modes review catches most often:
- It does not runThe most common defect: an environment that exists only on one laptop, never pinned.
- LeakageInformation from the evaluation set reaching the model through preprocessing fitted before the split. The result looks excellent and means nothing.
- The metric flattersA metric chosen after seeing the results, an imbalanced problem scored on accuracy, or a validation set reused until it became a training set.
- Undocumented dataA transformation applied by hand and never written down, so the pipeline cannot be repeated even by its author.
- Unstated limitsA deliverable that does not say what it did not test.
03
Tooling
| Tool | What it does here |
|---|---|
| GitHub | The deliverable itself. A repository with a history is what makes a submission reviewable: who changed what, when, and whether the result in the report corresponds to a commit that exists. |
| Jupyter | Where the analysis is done and shown. Also where most reproducibility failures start: a notebook whose cells were run out of order is a narrative, not a computation, so restart-and-run-all is the review standard. |
| Skrub | Preparing messy and heterogeneous tables for a model: assembling and encoding real-world data, including the dirty categorical columns that ordinary encoders handle badly. It removes the hand-written cleaning step that is usually the unrepeatable part of a pipeline. |
| scikit-learn | The baseline and the discipline: pipelines that fit preprocessing inside the cross-validation fold rather than before it, which is the mechanism that prevents leakage rather than a convention that asks people not to leak. |
| Optuna | Hyperparameter search as a recorded study rather than a manual sweep, so the search space and the trials are part of the artefact and the best trial can be pointed at. |
| MLflow | Recording runs with their parameters, metrics and artefacts, so that a claimed score is attached to the configuration that produced it. It answers the question a reviewer always has to ask: which run is this number from? |
| DVC | Versioning data and pipeline stages next to the code, so that a repository at a given commit identifies the data it was run on and which stages have to be re-executed after a change. |