When UK product teams ask us to shortlist architectures, we usually ask them to pause. Architecture debates are stimulating; labelling throughput decides whether a supervised approach can exist at all.
In a recent engagement we asked three annotators to label a sample of two hundred tickets under the draft guidelines. Disagreement rates above thirty percent on the minority class told us the taxonomy needed rewriting before any training run. That finding cost two days and saved a quarter of speculative modelling.
A practical rule: if you cannot name who will label the next five hundred examples, and at what weekly rate, you are not ready to compare gradient boosting with a neural net. Fix the definition of the target first.