Machine learning models contain two fundamental types of parameters: model parameters, which are learned during training via backpropagation or gradient descent, and hyperparameters, which are manually set before training begins. Hyperparameters directly control the learning process and model architecture. Common examples include learning rate, regularization strength (L1/L2), number of hidden layers, batch size, kernel size in CNNs, and tree depth in random forests.
Unlike model parameters, there is no closed-form mathematical solution for determining optimal hyperparameter values; they must be discovered empirically through systematic experimentation. Without proper tuning, even sophisticated models underperform dramatically. For instance, a neural network trained with a learning rate of 1.0 may diverge entirely, while a learning rate of 0.00001 may converge so slowly that training becomes impractical.
Grid Search is a brute-force yet effective hyperparameter optimization method that exhaustively evaluates every combination of predefined hyperparameter values. For each combination, a separate model is trained and validated, and the combination that produces the best validation performance is selected as the winner. This foundational technique remains widely used in production because it is interpretable, parallelizable, and guarantees finding the best combination within the defined search space, even though more sophisticated alternatives such as Random Search, Bayesian Optimization, and Hyperband exist for larger parameter spaces.
Analogy🏏Cricket
🏏 Think of it like cricket: In a cricket innings, Virat Kohli comes to bat and must decide his strategy—whether he'll play as an aggressive opener (like Rohit Sharma's powerplay style with big strokes) or as a stable middle-order anchor. His role, the type of deliveries he faces (fast bowlers vs. spin bowlers), and his run-scoring approach (boundaries vs. singles and doubles) are predetermined before he even steps into the crease. Similarly, when you create a variable in Python, you're assigning a 'role' to a memory location, specifying what 'type of data' it will hold (integer runs, string player names, boolean wicket status), and defining what 'operations' are valid on it. Just as a batsman cannot execute a reverse-sweep against a fast bowler at 145 km/h with the same technique he'd use against a spinner, a variable holding a string cannot perform arithmetic operations—you must first 'convert' or handle the type correctly. The cricket scorecard is the complete structure: each player has a name (string), a runs scored (integer), a balls faced (integer), and a dismissal status (boolean/string). Each of these data types has specific valid operations—you can add runs together, concatenate names for commentary, but you cannot add a player's name to their runs without explicit conversion, just as you cannot add a batsman's jersey number to his strike rate without understanding they represent different measurements.
🏏 Showing the Cricket analogy — a Cricket version isn’t available for this concept yet.