What is NanoFlow's scheduling policy? Is it the same as Sarathi-Server? How can we ensure that in each scheduling round, the number of decode tokens and prefill tokens remains fixed, so that the pipeline configuration found by auto-search can be properly utilized?
What is NanoFlow's scheduling policy? Is it the same as Sarathi-Server? How can we ensure that in each scheduling round, the number of decode tokens and prefill tokens remains fixed, so that the pipeline configuration found by auto-search can be properly utilized?