RabbieBurns Posted September 28, 2009 Posted September 28, 2009 anyone good with python? Im looking to get a python script that can parse an RSS feed and download the file that the RSS points to.. does anyone have a working example of this I might be able to use?
RabbieBurns Posted September 28, 2009 Author Posted September 28, 2009 ok ive been having a play with feedparser I can get the whole rss to list all the items, but im not sure how to iterate through them looking for the items i want, nor how to download the link associated to any matching items..
dhicks Posted September 29, 2009 Posted September 29, 2009 I can get the whole rss to list all the items, but im not sure how to iterate through them looking for the items i want, nor how to download the link associated to any matching items.. I just had a quick look at Universal Feed Parser. Seems easy enough - don't you want to just do a for-each loop over the entries list?: import feedparser d = feedparser.parse("http://feedparser.org/docs/examples/atom10.xml") for each entry in d['entries']: print entry -- David Hicks
RabbieBurns Posted September 29, 2009 Author Posted September 29, 2009 yeh that feedparser was what i meant id tried in my 2nd post.. what Im wanting to do is loop through the list, and compare the items to a pre-existing search criteria, and then download the links from any matching.. Im not really sure how to do anything to be honest so just experimenting with it all just now...
dhicks Posted September 29, 2009 Posted September 29, 2009 what Im wanting to do is loop through the list Which the list? The list of entries in the RSS file? import feedparser d = feedparser.parse("http://feedparser.org/docs/examples/atom10.xml") for each entry in d['entries']: Your code goes here. and compare the items to a pre-existing search criteria First, run the code above with the "print" statement in it, see what fields each "entry" dictionary item has. Decide which field you want to check, then you'll be able to do something along the lines of: import feedparser d = feedparser.parse("http://feedparser.org/docs/examples/atom10.xml") for each entry in d['entries']: if entry['name'] == "Bananas": print "We have bananas!" else: print "Sorry, we have no bananas." and then download the links from any matching There's bound to be a library for downloading resources from a given URL. Try searching the Python documentation with Google. -- David Hicks
RabbieBurns Posted September 29, 2009 Author Posted September 29, 2009 thanks, gives me a basis for thought processes.. say for example the rss is BBC News | News Front Page | World Edition then from a previous process (which Ive already got) the array of things to search for was, say "cricket" I would want it to just downlaod the .stm page that the rss points to.. but i think ive already got a routine to do the actual downloading part.. ach i dunno.,. Im not a coder whatsoever, I just muck about with others codes till i get something that resembles something that works.. ending up with a really crude effort
dhicks Posted September 29, 2009 Posted September 29, 2009 (edited) then from a previous process (which Ive already got) the array of things to search for was, say "cricket". I would want it to just downlaod the .stm page that the rss points to. Sorry, what part of what is "cricket"? A string you're looking for inside something, or the name of an array you've already read in from somewhere? The RSS file you link to would seem to have fields called title, description, link, guid, pubDate, category and media. Therefore, I'm guessing you want something along the lines of: import feedparser d = feedparser.parse("http://feedparser.org/docs/examples/atom10.xml") for entry in d['entries']: if entry['title'] == "cricket": goGetTheURL(entry['link']) -- David Hicks Edited September 29, 2009 by dhicks No "each" needed in for loop.
MrCurious Posted September 29, 2009 Posted September 29, 2009 Wow, timely thread. I'm doing the same thing. I'm getting stuck on the 'entry' iterator, unfortunately. >>> for each entry in d['entries'] File "", line 1 for each entry in d['entries'] ^ SyntaxError: invalid syntax >>> Any thoughts?
MrCurious Posted September 29, 2009 Posted September 29, 2009 argh, the shame. n/m! Helps if I write the for request in the method appropriate for the language. No "for each" statement
dhicks Posted September 29, 2009 Posted September 29, 2009 No "for each" statement Oh. Bother. ... -- David Hicks
dhicks Posted September 29, 2009 Posted September 29, 2009 (edited) Yep, just tested the following: import feedparser d = feedparser.parse("http://newsrss.bbc.co.uk/rss/newsonline_world_edition/front_page/rss.xml") for entry in d['entries']: if entry['title'] == "cricket": goGetTheURL(entry['link']) It goes and downloads an RSS file from the BBC website, checks each item listed and, if its title exactly matches "cricket", calls goGetTheURL with the entry's link. You'd probably want to check the title against a regular expression or something, an exact string match probably isn't what you want. RabbieBurns, MrCurious: Does that cover what you were trying to do? Post more details of what you're trying to do if you want. -- David Hicks Edited September 29, 2009 by dhicks
Recommended Posts
Create an account or sign in to comment
You need to be a member in order to leave a comment
Create an account
Sign up for a new account in our community. It's easy!
Register a new accountSign in
Already have an account? Sign in here.
Sign In Now