Repository navigation
Undeterministic behaviour?聽#2
Description
Activity
Somehow, TestCoCa always report 75% on this single test:
1025 1025But I do not see how it is possible to generate inputs that force a certain path, because the only inputs are the 2 lengths of the
string: the actual execution path will depend on the garbage that the memory block contains.The allocated memory is not initialized and it is then read in
cstrcspn. That is undefined behavior. TestCoCa executes the program natively. So there is very likely some random values (garbage) in the allocated memories. It is thus possible to evaluate branchings differently while processing both strings. The coverage then depends on the garbage (beside the two input numbers).Does this not make the results unreproducible?
The coverage result from TestCoCa may vary from run to run for the same two input numbers. The problem is not in TestCoCa. The problem is in the benchmark.
Does this not benefit test input generation tools that generate 1000s of random inputs?
Yes, I'd say so.
Would it not be better if the benchmark initialised the strings by reading characters up to its input length as verifier inputs in a loop?
Definitely. The current version of the benchmark is simply wrong. It should not be included in the competition.
How can I access the tests generated by KLEE
I do not know. I am also not right person to ask. I'd ask Dirk.
Somehow, TestCoCa always report 75% on this single test: 1025 1025
For such long strings the garbage inside them may have a property to cover the same set of branchings. I'd also say that for short stings (like 3 and 3) it is more likely the results would vary from one execution to another.
But I do not see how it is possible to generate inputs that force a certain path, because the only inputs are the 2 lengths of the
string: the actual execution path will depend on the garbage that the memory block contains.The allocated memory is not initialized and it is then read in
cstrcspn. That is undefined behavior. TestCoCa executes the program natively. So there is very likely some random values (garbage) in the allocated memories. It is thus possible to evaluate branchings differently while processing both strings. The coverage then depends on the garbage (beside the two input numbers).Does this not make the results unreproducible?
The coverage result from TestCoCa may vary from run to run for the same two input numbers. The problem is not in TestCoCa. The problem is in the benchmark.
Does this not benefit test input generation tools that generate 1000s of random inputs?
Yes, I'd say so.
Would it not be better if the benchmark initialised the strings by reading characters up to its input length as verifier inputs in a loop?
Definitely. The current version of the benchmark is simply wrong. It should not be included in the competition.
I agree with you, however, UB and reading uninitialised memory locations are allowed in Test-Comp ([Zulip](#Test-Comp > Undefined Behaviour @ 馃挰)). I do not think this is wise, but there you go.
I agree this is not a TestCoCa issue; I was just curious as to why the coverage was different from TestCov and wondered whether this was systematic or non-deterministic. But this is an academic discussion since the behaviour can be non-deterministic.
This is something to look into, although it may not be a TestCoCa issue.
As reported in [Zulip Thread](#SV-Benchmarks > Execution depending on garbage @ 馃挰):
I am looking at https://test-comp.sosy-lab.org/2026/results/sv-benchmarks/c/termination-crafted-lit/cstrcspn.c for the first time.
But I do not see how it is possible to generate inputs that force a certain path, because the only inputs are the 2 lengths of the string: the actual execution path will depend on the garbage that the memory block contains.
I have no doubt many similar examples exist.
Am I missing something?
Does this not make the results unreproducible?
Does this not benefit test input generation tools that generate 1000s of random inputs?
Would it not be better if the benchmark initialised the strings by reading characters up to its input length as verifier inputs in a loop?
Has this been discussed before?
The only tool that achieves 100% branch coverage is KLEE: although only TestCoCa reports 100%, Testcov x 4 versions all report 37.5%, which is suspicious.
How can I access the tests generated by KLEE: https://test-comp.sosy-lab.org/2026/results/results-verified/klee.2026-01-06_13-12-15.results.Test-Comp26_coverage-branches.C.coverage-branches.Termination-MainControlFlow.xml.bz2.fixed.xml.bz2.table.html#/table?filter=id_any(value(cstrcspn)) for curiosity?