Jump to content
EduGeek EdSec 2026 is Go! 27th Oct in Derby! Join us for a day of EdTech security focused talks, networking, and an evening social ×

Recommended Posts

Posted
anyone good with python? Im looking to get a python script that can parse an RSS feed and download the file that the RSS points to.. does anyone have a working example of this I might be able to use?
Posted

ok ive been having a play with feedparser

 

I can get the whole rss to list all the items, but im not sure how to iterate through them looking for the items i want, nor how to download the link associated to any matching items..

Posted
I can get the whole rss to list all the items, but im not sure how to iterate through them looking for the items i want, nor how to download the link associated to any matching items..

 

I just had a quick look at Universal Feed Parser. Seems easy enough - don't you want to just do a for-each loop over the entries list?:

 

import feedparser
d = feedparser.parse("http://feedparser.org/docs/examples/atom10.xml")
for each entry in d['entries']:
   print entry

 

--

David Hicks

Posted

yeh that feedparser was what i meant id tried in my 2nd post..

 

what Im wanting to do is loop through the list, and compare the items to a pre-existing search criteria, and then download the links from any matching..

 

Im not really sure how to do anything to be honest so just experimenting with it all just now...

Posted
what Im wanting to do is loop through the list

 

Which the list? The list of entries in the RSS file?

 

import feedparser
d = feedparser.parse("http://feedparser.org/docs/examples/atom10.xml")
for each entry in d['entries']:
   Your code goes here.

 

and compare the items to a pre-existing search criteria

 

First, run the code above with the "print" statement in it, see what fields each "entry" dictionary item has. Decide which field you want to check, then you'll be able to do something along the lines of:

 

import feedparser
d = feedparser.parse("http://feedparser.org/docs/examples/atom10.xml")
for each entry in d['entries']:
   if entry['name'] == "Bananas":
       print "We have bananas!"
   else:
       print "Sorry, we have no bananas."

 

and then download the links from any matching

 

There's bound to be a library for downloading resources from a given URL. Try searching the Python documentation with Google.

 

--

David Hicks

Posted

thanks, gives me a basis for thought processes..

 

say for example the rss is

 

BBC News | News Front Page | World Edition

 

then from a previous process (which Ive already got) the array of things to search for was, say "cricket"

 

I would want it to just downlaod the .stm page that the rss points to..

 

but i think ive already got a routine to do the actual downloading part..

 

ach i dunno.,. Im not a coder whatsoever, I just muck about with others codes till i get something that resembles something that works.. ending up with a really crude effort :)

Posted (edited)
then from a previous process (which Ive already got) the array of things to search for was, say "cricket". I would want it to just downlaod the .stm page that the rss points to.

 

Sorry, what part of what is "cricket"? A string you're looking for inside something, or the name of an array you've already read in from somewhere?

 

The RSS file you link to would seem to have fields called title, description, link, guid, pubDate, category and media. Therefore, I'm guessing you want something along the lines of:

 

import feedparser
d = feedparser.parse("http://feedparser.org/docs/examples/atom10.xml")
for entry in d['entries']:
   if entry['title'] == "cricket":
       goGetTheURL(entry['link'])

 

--

David Hicks

Edited by dhicks
No "each" needed in for loop.
Posted

Wow, timely thread.

 

I'm doing the same thing. I'm getting stuck on the 'entry' iterator, unfortunately.

 

>>> for each entry in d['entries']

File "", line 1

for each entry in d['entries']

^

SyntaxError: invalid syntax

>>>

 

 

Any thoughts?:ear:

Posted (edited)

Yep, just tested the following:

 

import feedparser
d = feedparser.parse("http://newsrss.bbc.co.uk/rss/newsonline_world_edition/front_page/rss.xml")
for entry in d['entries']:
if entry['title'] == "cricket":
	goGetTheURL(entry['link'])

 

It goes and downloads an RSS file from the BBC website, checks each item listed and, if its title exactly matches "cricket", calls goGetTheURL with the entry's link. You'd probably want to check the title against a regular expression or something, an exact string match probably isn't what you want.

 

RabbieBurns, MrCurious: Does that cover what you were trying to do? Post more details of what you're trying to do if you want.

 

--

David Hicks

Edited by dhicks

Create an account or sign in to comment

You need to be a member in order to leave a comment

Create an account

Sign up for a new account in our community. It's easy!

Register a new account

Sign in

Already have an account? Sign in here.

Sign In Now



×
×
  • Create New...