Skip to content

Undeterministic behaviour?聽#2

Description

@echancrure

This is something to look into, although it may not be a TestCoCa issue.
As reported in [Zulip Thread](#SV-Benchmarks > Execution depending on garbage @ 馃挰):

I am looking at https://test-comp.sosy-lab.org/2026/results/sv-benchmarks/c/termination-crafted-lit/cstrcspn.c for the first time.

But I do not see how it is possible to generate inputs that force a certain path, because the only inputs are the 2 lengths of the string: the actual execution path will depend on the garbage that the memory block contains.

I have no doubt many similar examples exist.

Am I missing something?
Does this not make the results unreproducible?
Does this not benefit test input generation tools that generate 1000s of random inputs?
Would it not be better if the benchmark initialised the strings by reading characters up to its input length as verifier inputs in a loop?
Has this been discussed before?

The only tool that achieves 100% branch coverage is KLEE: although only TestCoCa reports 100%, Testcov x 4 versions all report 37.5%, which is suspicious.

How can I access the tests generated by KLEE: https://test-comp.sosy-lab.org/2026/results/results-verified/klee.2026-01-06_13-12-15.results.Test-Comp26_coverage-branches.C.coverage-branches.Termination-MainControlFlow.xml.bz2.fixed.xml.bz2.table.html#/table?filter=id_any(value(cstrcspn)) for curiosity?

Activity

  1. echancrure commented on Mar 18, 2026

    @echancrure
    Author

    Somehow, TestCoCa always report 75% on this single test:

    1025 1025
  2. trtikm commented on Mar 20, 2026

    @trtikm
    Contributor

    But I do not see how it is possible to generate inputs that force a certain path, because the only inputs are the 2 lengths of the
    string: the actual execution path will depend on the garbage that the memory block contains.

    The allocated memory is not initialized and it is then read in cstrcspn. That is undefined behavior. TestCoCa executes the program natively. So there is very likely some random values (garbage) in the allocated memories. It is thus possible to evaluate branchings differently while processing both strings. The coverage then depends on the garbage (beside the two input numbers).

    Does this not make the results unreproducible?

    The coverage result from TestCoCa may vary from run to run for the same two input numbers. The problem is not in TestCoCa. The problem is in the benchmark.

    Does this not benefit test input generation tools that generate 1000s of random inputs?

    Yes, I'd say so.

    Would it not be better if the benchmark initialised the strings by reading characters up to its input length as verifier inputs in a loop?

    Definitely. The current version of the benchmark is simply wrong. It should not be included in the competition.

    How can I access the tests generated by KLEE

    I do not know. I am also not right person to ask. I'd ask Dirk.

  3. trtikm commented on Mar 20, 2026

    @trtikm
    Contributor

    Somehow, TestCoCa always report 75% on this single test: 1025 1025

    For such long strings the garbage inside them may have a property to cover the same set of branchings. I'd also say that for short stings (like 3 and 3) it is more likely the results would vary from one execution to another.

  4. echancrure commented on Mar 20, 2026

    @echancrure
    Author

    But I do not see how it is possible to generate inputs that force a certain path, because the only inputs are the 2 lengths of the
    string: the actual execution path will depend on the garbage that the memory block contains.

    The allocated memory is not initialized and it is then read in cstrcspn. That is undefined behavior. TestCoCa executes the program natively. So there is very likely some random values (garbage) in the allocated memories. It is thus possible to evaluate branchings differently while processing both strings. The coverage then depends on the garbage (beside the two input numbers).

    Does this not make the results unreproducible?

    The coverage result from TestCoCa may vary from run to run for the same two input numbers. The problem is not in TestCoCa. The problem is in the benchmark.

    Does this not benefit test input generation tools that generate 1000s of random inputs?

    Yes, I'd say so.

    Would it not be better if the benchmark initialised the strings by reading characters up to its input length as verifier inputs in a loop?

    Definitely. The current version of the benchmark is simply wrong. It should not be included in the competition.

    I agree with you, however, UB and reading uninitialised memory locations are allowed in Test-Comp ([Zulip](#Test-Comp > Undefined Behaviour @ 馃挰)). I do not think this is wise, but there you go.

    I agree this is not a TestCoCa issue; I was just curious as to why the coverage was different from TestCov and wondered whether this was systematic or non-deterministic. But this is an academic discussion since the behaviour can be non-deterministic.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions