Repository navigation
Parallelize CodeSplitter's not-exclusive CFA computation - #10396
rhuanhianc wants to merge 1 commit into
Conversation
computeNotExclusiveCfaForFragments traverses the run-asyncs of every other fragment for each fragment, so it is quadratic in fragment count. The iterations are independent, so run them on a thread pool. The serial loop is kept for when a dependency recorder is installed, since recording is order sensitive. JProgram.getTypeArray is now synchronized, being the only shared state the traversals create on demand. CodeSplitter goes from 181.4s to 71.5s on an application with 231 split points, with byte-identical output.
niloc132
left a comment
There was a problem hiding this comment.
I'm not sure about this - at the very least, tying it to localWorkers will take the sting off of it a bit, but we still risk actually running localWorkers^2 concurrent threads (and running out of memory by doing so much at once, or other issues from too much work scheduled).
For threaded workers, we could at least share a single ExecService with extra localWorker count (e.g. localWorkers - permutationCount threads haven't been used yet, let them be scheduled for this work, plus or minus how many permutations are paused waiting for their exclusive fragments). But when we use external processes for workers (e.g. ExternalPermutationWorkerFactory), we don't have as nice of a way to communicate about which process should get the extra threads...
Mechanically I don't much to complain about though, just want to consider this approach carefully before risking breaking anyone's CI.
| // Synchronized because CodeSplitter reaches this from concurrent CFA traversals. | ||
| public synchronized JArrayType getTypeArray(JType elementType) { |
There was a problem hiding this comment.
Instead, perhaps just call program.getTypeArray(JPrimitiveType.INT) once before branching so we guarantee the type exists. I think we could even drop that call entirely, since it only exists for an assert (and once it matches a == check, we just use the same instance).
computeNotExclusiveCfaForFragmentsis quadratic in the number of exclusive fragments: for each fragment it traverses the run-asyncs of every other fragment. Applications with many split points pay heavily — with 231 split points this loop alone is 25.6% of the permutation's CPU, and CodeSplitter overall is 181s of an 11-minute compile.The iterations are independent.
ControlFlowAnalyzer's copy constructor deep-copies every mutable set, each traversal writes only to its own analyzer, andControlFlowAnalyzeris purely analytical — it contains no setters on AST nodes. This runs them on a fixed thread pool.Two constraints are respected:
-compileReportthe original serial loop runs unchanged.JProgram.getTypeArraylazily creates array types in a plain HashMap and is reachable from the traversal, so it is nowsynchronized. It is the only shared mutable state on this path;getAllArrayTypes, which already sorts to avoid nondeterminism, is not reachable fromtraverseFromRunAsync.Happy to derive the pool size from
-localWorkersinstead ofavailableProcessors()if you'd prefer to avoid oversubscription when permutations are already compiled in parallel.CodeSplitter drops from 38.4s to 17.0s with 122 split points and from 181.4s to 71.5s with 231. Generated JavaScript is byte-for-byte identical across 874 output files from four applications. ant -Dtarget=test dev passes: 1936 tests, 0 failures.
Fixes #10395