The run size preview for airt.multilingual leaves out the language axis, so it undercounts by 3.5x on a default run and more with more languages.
Multilingual doesn't override _estimate_run_size_async, so it gets the base "seed groups x techniques + baseline" count. The real run builds one translation slice per language plus a random-translation slice, each over every technique and seed group.
Repro with 5 seed groups and the default technique, estimate vs the initialized run:
default (5 random languages): estimate total=10 | initialized units=35
languages=[8 languages]: estimate total=10 | initialized units=50
num_languages, languages and translation_strategies don't change the estimate at all.
The run size preview for airt.multilingual leaves out the language axis, so it undercounts by 3.5x on a default run and more with more languages.
Multilingual doesn't override
_estimate_run_size_async, so it gets the base "seed groups x techniques + baseline" count. The real run builds one translation slice per language plus a random-translation slice, each over every technique and seed group.Repro with 5 seed groups and the default technique, estimate vs the initialized run:
num_languages,languagesandtranslation_strategiesdon't change the estimate at all.