Skip to content

Commit e1eba80

Browse files
authored
Update README.md
1 parent 27189b3 commit e1eba80

File tree

1 file changed

+4
-0
lines changed

1 file changed

+4
-0
lines changed

README.md

Lines changed: 4 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -7,3 +7,7 @@ Notebooks run based on local path (it should be the github repo one for local ru
77

88
### Limits of statistical features
99
By using only informations about the number of the words, their length, and similar things about the sentences or the text as a whole does not make a decent result: RMSE is about 0.81 in most cases, even using a minimal amount of augmentation. However it could be useful to further boost more advanced models. Time will tell :)
10+
11+
### Proposed approaches
12+
13+
* From a rapid look at the data, it seems that the most readable texts talk about well known (by a kid) things in a simple way (for example: dinosaurs), while the least readable one are about highly techinical subjects (such as metalworking), with a lot of technical words in the text (which of course will be hard to read for a kid). This suggests two approaches: the first one is to select the argument(s) of the text as an extra training variable (maybe spacy + W2V), the second one is to estimate the frequency of words in a large corpora of WELL BALANCED text and look for those words that are higly specilized and unusual (maybe accounting also for the intrinsic "intensione (vedi 'Il software del linguaggio' di R. Simone)

0 commit comments

Comments
 (0)