Jump to content

Recommended Posts

Posted

I've just realised that our school filter does not block Andrew Tate's own website for pupils, nor does it prevent pupils from signing up. (cobratate.com).

 

From the marketing fluff, I wouldn't be expecting to have to block this as manual url block myself.

 

Does yours?

  • Thanks 3
Posted
Thanks for posting this. Wasn't blocked here.

 

Some other URLs being used from that site:

 

cobratate.com

jointherealworld.com

thewarroom.ag

rumble.com/c/TateSpeech

gettr.com/user/cobratate

cobratatemembers.com

Frustratingly, every one of these except for Rumble was allowed by default in lightspeed...

 

I found students searching him around six months ago and blocked a number of search terms and sites. There's a number of "hustler academy" sites to block too. Shame the search filter doesn't do much as they can just put a space or misspell anything to get around it.

Posted (edited)

My filter does have a Hate Speech category, but doesn't include Tate.

 

None of the sites listed above - some of which should have been categorised as social media or streaming at least just sailed straight into the browser.

Edited by sigma
Posted
My filter does have a Hate Speech category, but doesn't include Tate.

 

None of the sites listed above - some of which should have been categorised as social media or streaming at least just sailed straight into the browser.

 

That is all kinds of worrying. I'd think most if not all of us would agree that Tate and the like are extremely dangerous for young males to be exposed to.

Posted
That is all kinds of worrying. I'd think most if not all of us would agree that Tate and the like are extremely dangerous for young males to be exposed to.

 

Not as dangerous as they are to women.

 

Still, Uncle Elon says he's OK.

  • 5 months later...
Posted
Filtering doesn't think, it's based on the content of the page, or someone reporting it

 

I entirely agree. The "AI-based categorisation tool" seems to be largely vapourware. I reckon its a big url database manually updated by idiots like me reporting sites.

Posted

Can still be AI based, that's just fancy statistics. "here's 10000 pages we manually said are Bad, find out what's similar about them and set other sites that are similar to Bad"

 

It doesn't scan the news and whois databases

Posted
I entirely agree. The "AI-based categorisation tool" seems to be largely vapourware. I reckon its a big url database manually updated by idiots like me reporting sites.

 

I'm pretty confident that AI/ML can easily be excellent at categorizing previously unseen pages, however it can't current do it fast enough to be real time without an enormous power requirements and at the scale required for a trust wide / regional system. For a single site, sling a couple of Nvidia Grace Hoppers in a rack in each school and you'd be good.

Posted (edited)
you only need to do it once per page/site and then distribute the results to all clients

Except most pages these days are dynamic, so that "once per url" strategy is not as effective as it once was.

 

You're not retraining on every site, you're using the previously trained model.

 

There'd still be a delay as it evaluated the page, modern consumer cpus have their tensor / neural engine cores and would probably be fine at this, but a network level system would need a serious amount of grunt to process 1Gb/s+ of web traffic.

 

https://data.world/crowdflower/url-categorization someone's already gathered data for you as well

Cool. I'll just knock one up. BRB 😂

Edited by psydii
Posted
the more fine grained filtering must be done at the client level due to the dynamic nature of websites, docs.google.com can be anything
Posted

AI can certainly pick up things like a page full of porn, but keyword filters work well for that sort of stuff too and need less resources. Where keyword filters struggle, AI will also struggle, because there are mainly 2 failure modes:

 

1. Some content just doesn't contain anything that looks concerning, when taken in isolation. A lot of the Andrew Tate stuff is a good example - you need to understand who Tate is, his background, what the problems are, etc. in order to know that it's concerning content. Current AI technologies are usually run through an extensive training cycle once, and then used to analyse content in isolation. Keeping their training up to date so that they can take this kind of background context into account is really hard.

 

2. Humans want to believe that things can be divided into nice simple categories, but that simply isn't true - its often extremely difficult to decide how to categorise a web page. We have written categorisation criteria for (human) analysts to follow when manually recategorising websites, but still have regular discussions internally on how to categorise specific pieces of content. Different members of staff often have conflicting opinions - this stuff is not at all clear cut. The results of those discussions not only lead to a decision being made on specific content, but feeds back into the categorisation criteria that will be used in the future.

For example, what specifically should a "Weapons" category block? If you block all websites selling knives then you're blocking a really wide range of stuff (keyring penknives, craft knives, Stanley knives, kitchenware, right up to the stuff someone might carry into a fight). Wikipedia pages describing commercially available hand guns may seem like fair game, but what about Wikipedia pages documenting the weapons used in the world wars? Legitimate gun clubs / firing ranges? Info about / sale of digital items (e.g. video game weapons)? Books / movies / TV programmes which contain weapons? Arguably all of that is "weapons" but blocking it all would probably be harmful to teaching.

 

Similar problems for a "Violence" category - News articles? Historical accounts? Fictional stories? The Bible?.

 

Porn is the classic "I know it when I see it", and from experience different schools have very very different opinions on what they think is "porn", with some even considering the front covers of some mainstream papers that you would see if you ventured into your local Tesco to be "porn" and therefore blockable. What about historical paintings? What sets a classical nude painting apart from a modern nude cartoon? etc.

 

These are all human problems, and are hard for humans to grapple with. Definitely not something that an AI is qualified to make decisions on.

 

(For what it's worth, we have a separate "Tate Brothers" category, because some schools want to block it, others want to monitor it and other's don't seem to care!).

  • Thanks 4

Create an account or sign in to comment

You need to be a member in order to leave a comment

Create an account

Sign up for a new account in our community. It's easy!

Register a new account

Sign in

Already have an account? Sign in here.

Sign In Now



×
×
  • Create New...