Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

Can IA be mirrored like Wikipedia?


It seems like there might be all kinds of copyright problems when mirroring IA. With Wikipedia, everything is licensed under Creative Commons. But IA is a copy of all kinds of websites with all kinds of licenses. It's a big effort for them simply to remove content that companies rightfully request be removed.


Much (all?) of IA content already has metadata for licensing, e.g. out-of-copyright books. It should be legally ok to mirror the CC and public domain content, but I don't know if there's an easy way to extract a dump or torrent.


Check out robots.txt -- most companies use it (in this case) because it's easier than dmca takedowns.




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: