**Background:** Genomic selection (GS) can accelerate plant breeding by predicting phenotypic values from genotypic data, but its practical implementation is limited by the high cost of multi-environment trials. Sparse testing—where only a subset of lines is evaluated in each environment—can reduce costs while maintaining prediction accuracy. Previous studies used random allocation without optimization and only uni-trait models. This study evaluated four allocation methods under both uni-trait and multi-trait frameworks to determine optimal sparse testing strategies.
**Methods:** Two CIMMYT datasets were used: a maize dataset with 484 lines evaluated in three environments (drought stress, low nitrogen, well-watered) for grain yield and plant height, and a wheat dataset with 4,464–4,536 lines evaluated in four environments (B2IR, B5IR, BDRT, BLHT) for grain yield, days to germination, days to heading, and plant height. Smaller subsets of 250 lines were also created from each dataset. Four sparse testing methods were compared: M1 (fraction of lines in all locations), M2 (fraction with some shared lines across locations), M3 (random allocation under incomplete locations), and M4 (allocation using balanced incomplete block design principles). A genomic prediction model including genotype × environment interaction was used. Cross-validation with 10 partitions was performed at five testing proportions (15%, 25%, 50%, 75%, 85%). Prediction accuracy was measured using Average Pearson's Correlation (APC) and Normalized Root Mean Squared Error (NRMSE).
**Key Results:** Across all datasets, multi-trait models consistently outperformed uni-trait models. For the complete maize dataset under 15% testing with uni-trait models, M4 outperformed M1, M2, and M3 by 12.89%, 7.08%, and 7.33% respectively in APC. Under 85% testing uni-trait, M4 outperformed M1 and M2 by 1.6% and 4.9%. In NRMSE for maize at 85% testing, M4 outperformed M1 and M2 by 5% and 60.6% (multi-trait) and by 3.7% and 30.5% (uni-trait). For the complete wheat dataset under 15% testing uni-trait, M4 outperformed M1 and M2 by 24.13% and 24.08% in APC. Across all datasets aggregated, at 15% testing uni-trait, M3 and M4 outperformed M1 and M2 by 8.6% and 3.9% in APC. In NRMSE across datasets at 15% testing uni-trait, M4 outperformed M1 and M2 by 8.0% and 5.8%. M3 showed the most consistent and robust performance overall. The cost-benefit analysis showed that with a 50–50% training-testing split, breeders could increase lines under evaluation by 101.12% with the same fixed budget. Even at the extreme 15–85% training-testing split, new lines evaluated increased by at least 573%.
**Clinical Implications:** While this is a plant breeding study rather than a clinical trial, the findings have significant implications for agricultural productivity and food security. Sparse testing methods M3 and M4 enable breeders to evaluate substantially more candidate lines under fixed budgets without meaningful loss of prediction accuracy. Multi-trait models should be preferred over uni-trait models. The random allocation method (M3) is recommended for large datasets due to computational efficiency, while the balanced incomplete block design method (M4) is preferred when possible for its optimality properties. These strategies can accelerate genetic gain in staple crops like maize and wheat, contributing to global food security.