text mining brainstorm
there are a number of techniques that may do the job, but each of them one of more of the following:
- computationally expensive
- bring large dependencies
- not really designed for Italian lang 🇮🇹
As a consequence we (@giuliogcantone and me) tried to figure out what we can do, these are some of the proposals:
- sort of topic mining on contract objects compared to a manually annotated dictionary from domain experts ❌
- 0-shot classification (italian using roberta) on very specific critical terms that may let us think there's something suspicious with it ❌
- from
cpv exact description and compare against the objects over a number of similarity measures (Levenshtein, Jaccard etc) ✅
We are currently investigating the third solution, but we are really open to discuss the other two.
text mining brainstorm
there are a number of techniques that may do the job, but each of them one of more of the following:
As a consequence we (@giuliogcantone and me) tried to figure out what we can do, these are some of the proposals:
cpvexact description and compare against the objects over a number of similarity measures (Levenshtein, Jaccard etc) ✅We are currently investigating the third solution, but we are really open to discuss the other two.