November 17, 2008

Pinning down the fleeting Internet: Web crawler archives historical data for easy searching

(PHysOrg.com) -- The Internet contains vast amounts of information, much of it unorganized. But what you see online at any given moment is just a snapshot of the Web as a whole -- many pages change rapidly or disappear completely, and the old data gets lost forever.

"Your browser is really just a window into the Web as it exists today," said Eytan Adar, University of Washington computer science and engineering doctoral student. "When you search for something online, you're only getting today's results."

Now, Adar and his colleagues at UW and Adobe Systems Inc. are grabbing hold of the fleeting Web and storing historical sites that users can easily search using an intuitive application called Zoetrope.

"There are so many ways of finding and manipulating and visualizing data on what we call 'the today Web' that it's kind of amazing that there's no way to do anything similar to the ephemeral Web," said Dan Weld, a UW computer science and engineering professor who also worked on the application. One service, the Internet Archive, has been capturing old versions of Web sites for years, but the records for the stored sites are inconsistent, Weld said. More importantly, there's no easy way to search the archive.

With Zoetrope, anyone will be able to use easy keyword searches to find archived Web information or look for patterns over time. The research was presented Oct. 22 by Mira Dontcheva, the system's co-creator and a recently graduated UW computer science and engineering doctoral student now at Adobe Systems Inc., at the ACM Symposium on User Interface Software and Technology in Monterey, Calif.

There are a variety of ways people might want to search the historical Internet. For example, to find a history of traffic patterns in the Seattle area, you'd have to sort through lengthy PDF files from the state Department of Transportation, Adar said. With Zoetrope, you could easily view past versions of any traffic Web site, and getting more specific, search for drive-times on Interstate 90 at 6 p.m. on rainy Fridays. Zoetrope can also capture and help analyze information that might otherwise not be available anywhere.

Sports fanatics could use the program to check historical rankings of their favorite teams or players, information that currently may not be easy to find. The application can do more than just simple keyword searches, Adar said. It also can be used to analyze historical data or link information from different sites. For example, Adar wondered whether air pollution conditions could affect the performance of Olympic athletes, so he used Zoetrope to find daily records of pollution levels in Beijing and the number of world records broken in the 2008 Olympics on each day, and looked to see whether fewer records were broken on days with high pollution levels.

"Zoetrope is aimed at the casual researcher," Weld said. "It's really for anyone who has a question."

Zoetrope could eventually be built in to any other Web browser, Adar said. If you just want to browse the past versions of a given site, you drag a slider backwards to see older and older versions. Alternatively, you can draw a box around just one part of the site, if you're interested in, say, the lead story on CNN.com but don't care about the rest of the page. These boxes can be filtered by keyword searches or date, so you could look only for lead stories featuring Hollywood actors or stories that ran on Fridays.

Users can view historical data by moving the slider, but more sophisticated analyses are available as well. If you're looking at something numerical, such as gas prices over time, the program can draw graphs for you. Or you can pull out images from specific times, such as traffic pictures, and compare them all side by side. These kinds of visualizations can be further organized in a timeline or by clustering -- Zoetrope can make an image comparing traffic patterns on sunny days versus cloudy days, for example.

Right now, Zoetrope saves a new version of approximately 1,000 different sites every hour, Adar said. It's been running for four months, so records go no further than that, but Adar hopes to eventually incorporate information from the Internet Archive's nearly 14 years of records into the program.

He wants to figure out how to scale the program up from 1,000 Web pages to all pages in existence, and has run studies to figure how often each page would need to be captured. For example, a traffic site or stock-watching page would need versions saved much more often than every hour, but there are many unchanging pages that could be archived less frequently. Eventually, Zoetrope could automatically figure out how often to capture a page based on how frequently it changes, Adar said.

"This is really a new way to think about storing information on the Web," he said.

The researchers hope to release Zoetrope free, and say it may be available as early as next summer.

Provided by University of Washington

Citation: Pinning down the fleeting Internet: Web crawler archives historical data for easy searching (2008, November 17) retrieved 19 September 2024 from https://phys.org/news/2008-11-pinning-fleeting-internet-web-crawler.html

This document is subject to copyright. Apart from any fair dealing for the purpose of private study or research, no part may be reproduced without the written permission. The content is provided for information purposes only.

Explore further

Wolves reintroduced to Isle Royale temporarily affect other carnivores, humans have influence as well

0 shares

Feedback to editors

Scientists unearth key clues to cuisine of resident killer whale populations

7 minutes ago

NASA develops process to create very accurate eclipse maps

12 minutes ago

Fossil site in Massachusetts reveals 320-million-year-old ecosystem

22 minutes ago

Volcanoes may help reveal interior heat on Jupiter moon

31 minutes ago

New study uncovers unexpected interaction between Mars and the solar wind

48 minutes ago

Learning mindset could be key to addressing medical students' alarming burnout

51 minutes ago

Mussel-inspired adhesive comes unglued on command

52 minutes ago

10,000-year-old human DNA provides insights into South African population history

52 minutes ago

The mystery of human wrinkles: What do the cells say?

59 minutes ago

AI model can reveal the structures of crystalline materials

59 minutes ago

Load comments (0)

Pinning down the fleeting Internet: Web crawler archives historical data for easy searching

Scientists unearth key clues to cuisine of resident killer whale populations

NASA develops process to create very accurate eclipse maps

Fossil site in Massachusetts reveals 320-million-year-old ecosystem

Volcanoes may help reveal interior heat on Jupiter moon

New study uncovers unexpected interaction between Mars and the solar wind

Learning mindset could be key to addressing medical students' alarming burnout

Mussel-inspired adhesive comes unglued on command

10,000-year-old human DNA provides insights into South African population history

The mystery of human wrinkles: What do the cells say?

AI model can reveal the structures of crystalline materials

Relevant PhysicsForums posts

Container shrinks at certain screen widths (CSS)

Unsolvable python code bug? (finding the difference between two input strings)

User-Defined Functions in Sql Server SSMS

Can Fortran 77 Code Be Used to Debug Python Code for Solving ODEs Using Radau5?

Help solving a geometrical matching issue with Graph Neural Networks

Zipping identical iterables

Wolves reintroduced to Isle Royale temporarily affect other carnivores, humans have influence as well

Plankton researchers urge their colleagues to mix it up

Crime-terror nexus obstructs global fight against illicit drugs

The fascinating sex lives of insects

Model shows how plankton survive in a turbulent world

Interactive map shows future climate of your city based on emissions scenarios

Hyphens in paper titles harm citation counts and journal impact factors

A big step toward the practical application of 3-D holography with high-performance computers

Combining multiple CCTV images could help catch suspects

Applying deep learning to motion capture with DeepLabCut

Training artificial intelligence with artificial X-rays

New model for large-scale 3-D facial recognition

Medical Xpress

Tech Xplore

Science X

Pinning down the fleeting Internet: Web crawler archives historical data for easy searching

Scientists unearth key clues to cuisine of resident killer whale populations

NASA develops process to create very accurate eclipse maps

Fossil site in Massachusetts reveals 320-million-year-old ecosystem

Volcanoes may help reveal interior heat on Jupiter moon

New study uncovers unexpected interaction between Mars and the solar wind

Learning mindset could be key to addressing medical students' alarming burnout

Mussel-inspired adhesive comes unglued on command

10,000-year-old human DNA provides insights into South African population history

The mystery of human wrinkles: What do the cells say?

AI model can reveal the structures of crystalline materials

Relevant PhysicsForums posts

Related Stories

Wolves reintroduced to Isle Royale temporarily affect other carnivores, humans have influence as well

Plankton researchers urge their colleagues to mix it up

Crime-terror nexus obstructs global fight against illicit drugs

The fascinating sex lives of insects

Model shows how plankton survive in a turbulent world

Interactive map shows future climate of your city based on emissions scenarios

Recommended for you

Hyphens in paper titles harm citation counts and journal impact factors

A big step toward the practical application of 3-D holography with high-performance computers

Combining multiple CCTV images could help catch suspects

Applying deep learning to motion capture with DeepLabCut

Training artificial intelligence with artificial X-rays

New model for large-scale 3-D facial recognition

Newsletter sign up

Donate and enjoy an ad-free experience