Volunteering / technical training and review

Volunteer role

AI for Impact

Technical Trainer, Deliverables Reviewer & Organizer

Role

Title
Technical Trainer, Deliverables Reviewer & Organizer
Type
Volunteer
Training
Preparing and delivering technical training to participants
Support
Technical support to teams during the event
Review
Analysis and review of technical deliverables
Organisation
Hackathon organisation

What is reviewed for

Reproducibility
Can someone else run this and get the same thing?
Data methodology
Where the data came from, what was done to it, and whether the split is honest
ML methodology
Whether the result is a result or an artefact of the evaluation
Clarity
Whether the deliverable states what it did and what it did not
01

The role

  • Technical trainingPreparing and delivering training on the tooling and the method: how to set a project up so that it can be handed to someone else, how to handle data, and how to evaluate a model without fooling yourself.
  • Participant supportTechnical support during the event itself, where the questions are concrete and the clock is running: an environment that will not build, a pipeline that leaks, a metric that looks too good.
  • Deliverable reviewReading what teams hand in and assessing it technically, which mostly means asking whether the number in the summary survives someone else rerunning the notebook.
  • OrganisationHackathon organisation.
02

What review actually catches

The failure modes review catches most often:

  • It does not runThe most common defect: an environment that exists only on one laptop, never pinned.
  • LeakageInformation from the evaluation set reaching the model through preprocessing fitted before the split. The result looks excellent and means nothing.
  • The metric flattersA metric chosen after seeing the results, an imbalanced problem scored on accuracy, or a validation set reused until it became a training set.
  • Undocumented dataA transformation applied by hand and never written down, so the pipeline cannot be repeated even by its author.
  • Unstated limitsA deliverable that does not say what it did not test.
03

Tooling

Tooling and its engineering role in the workflow
ToolWhat it does here
GitHubThe deliverable itself. A repository with a history is what makes a submission reviewable: who changed what, when, and whether the result in the report corresponds to a commit that exists.
JupyterWhere the analysis is done and shown. Also where most reproducibility failures start: a notebook whose cells were run out of order is a narrative, not a computation, so restart-and-run-all is the review standard.
SkrubPreparing messy and heterogeneous tables for a model: assembling and encoding real-world data, including the dirty categorical columns that ordinary encoders handle badly. It removes the hand-written cleaning step that is usually the unrepeatable part of a pipeline.
scikit-learnThe baseline and the discipline: pipelines that fit preprocessing inside the cross-validation fold rather than before it, which is the mechanism that prevents leakage rather than a convention that asks people not to leak.
OptunaHyperparameter search as a recorded study rather than a manual sweep, so the search space and the trials are part of the artefact and the best trial can be pointed at.
MLflowRecording runs with their parameters, metrics and artefacts, so that a claimed score is attached to the configuration that produced it. It answers the question a reviewer always has to ask: which run is this number from?
DVCVersioning data and pipeline stages next to the code, so that a repository at a given commit identifies the data it was run on and which stages have to be re-executed after a change.