Hypothesis: prompt optimization will not work too well for this task. Why?
(1) The skill in forecasting seems more "tacit", and less discrete rules or knowledge. The latter is where reflective prompt optimization (e.g. DSPy) shines.
(2) We see benefits from scaling training data (Figure 3), and due to context length limitations, this might not help GEPA-like techniques much yet.
It would be great if someone tests this out, as it's both somewhat low hanging, and also tests important hypotheses like the ones above^
Hypothesis: prompt optimization will not work too well for this task. Why?
(1) The skill in forecasting seems more "tacit", and less discrete rules or knowledge. The latter is where reflective prompt optimization (e.g. DSPy) shines.
(2) We see benefits from scaling training data (Figure 3), and due to context length limitations, this might not help GEPA-like techniques much yet.
It would be great if someone tests this out, as it's both somewhat low hanging, and also tests important hypotheses like the ones above^