-
Notifications
You must be signed in to change notification settings - Fork 456
timbertson/python-readability
Folders and files
Name | Name | Last commit message | Last commit date | |
---|---|---|---|---|
Repository files navigation
This code is under the Apache License 2.0. http://www.apache.org/licenses/LICENSE-2.0 This is a python port of a ruby port of arc90's readability project http://lab.arc90.com/experiments/readability/ Given a html document, it pulls out the main body text and cleans it up. Ruby port by starrhorne and iterationlabs Python port by gfxmonk This port uses BeautifulSoup for the HTML parsing. That means it can be a little slow, but will work on Google App Engine (unlike libxml-based libraries) **note**: I don't currently have any plans for using or improving this library, and it's far from perfect (slow, and almost certainly buggy). So if you do something cool with it or have a better tool that does the same job, please let me know and I can link to it from here. If you're looking for alternatives / forks, here's the list so far: - http://www.minvolai.com/blog/decruft-arc90s-readability-in-python/ - https://github.com/buriy/python-readability
About
[abandoned] python port of arc90's readability bookmarklet
Resources
Stars
Watchers
Forks
Releases
No releases published
Packages 0
No packages published