August 8, 2007

New search engine ranks tables by title, document content, text reference

Penn State researchers have developed a search engine-TableSeer-which not only can identify and extract tables from PDF documents but also can index and rank the search results using factors including the table's title, text references to the table and date of publication.

The engine's innovative ranking algorithm, TableRank, also can identify tables found in frequently cited documents and weigh that factor as well in the search results, said Prasenjit Mitra, an assistant professor in the Penn State College of Information Sciences and Technology (IST) and one of the lead researchers in the development of the search engine.

"TableSeer makes it easier for scientists and scholars to find and access the important information presented in tables, and as far as we know, is the first search engine for tables," Mitra said.

Tables are an important data resource for researchers. In a search of 10,000 documents from s and conferences, the researchers found that more than 70 percent of papers in chemistry, biology and computer science included tables. Furthermore, most of those documents had multiple tables.

But while some software can identify and extract tables from text, existing software cannot search for tables across documents. That means scientists and scholars must manually browse documents in order to find tables-a time-consuming and cumbersome process.

TableSeer automates that process and captures data not only within the table but also in tables' titles and footnotes. In addition, it enables column-name-based search so that a user can search for a particular column in a table.

In tests with documents from the Royal Society of Chemistry, TableSeer correctly identified and retrieved 93.5 percent of tables created in text-based formats, Mitra said.

Searching for tables has some unique challenges, as there is no standard table representation, so tables can appear in PDF, PowerPoint, HTML and Microsoft Word documents. The researchers chose to focus on PDF documents because of their growing popularity in digital libraries and because PDF documents had been overlooked in other table-search efforts.

"Tables can be made using a number of editor tools, and the techniques we are using in TableSeer should work with any text-based tool," said C. Lee Giles, professor of information sciences and technology and co-director of the IST Cyber-Infrastructure Lab where the research originated. "While we designed and developed TableSeer to facilitate searching of tables occurring in articles in the chemistry domain, it can be used in any domain where data is presented in tabular form including other scientific, technical, social and business areas."

The development of TableSeer is part of an open-source cyber-infrastructure project focusing on chemical document search for environmental chemistry and funded by the National Science Foundation. The grant awarded to the Penn State Department of Chemistry aims to enable automatic data analysis.

"Searching and extracting information from data tables is an essential component of data analysis in environmental science, where many research groups publish large amounts of kinetic data describing chemical changes in the environment," said Karl Mueller, professor of chemistry and principal investigator for the NSF grant.

"As we approach multidisciplinary problems within the Penn State Center for Environmental Kinetics Analysis, our students spend many days hunting down and compiling large amount of data from tables. The TableSeer tools will definitely increase the efficiency of this process and allow more time to be spent on creative scientific analysis," he added.

TableSeer can be tested online (see chemxseer.ist.psu.edu). The source code will be made available near the completion of the project, the researchers said.

In the meantime, research is ongoing to improve the ranking algorithm by adding additional features. The researchers also are working on a search engine that can identify, extract and rank figures found in documents, as figures are another important device for disseminating data and findings in the natural sciences.

Source: Penn State

Citation: New search engine ranks tables by title, document content, text reference (2007, August 8) retrieved 7 August 2024 from https://phys.org/news/2007-08-tables-title-document-content-text.html

This document is subject to copyright. Apart from any fair dealing for the purpose of private study or research, no part may be reproduced without the written permission. The content is provided for information purposes only.

Explore further

New calculation approach allows more accurate predictions of how atoms ionize when impacted by high-energy electrons

0 shares

Feedback to editors

Modern aircraft emit less carbon than older aircraft, but their contrails may do more environmental harm

1 hour ago

Scientists equip Australian sea lions with cameras to explore previously unmapped ocean habitats

2 hours ago

Fishing disrupts squaretail grouper mating behavior, study finds

4 hours ago

Domestication causes smaller brain size in dogs than in the wolf: Study challenges notion

7 hours ago

Tundra vegetation to grow taller, greener through 2100, study finds

10 hours ago

Living with a killer: How an unlikely mantis shrimp-clam association violates a biological principle

11 hours ago

Bouncing helps people move in sync during dance, study shows

11 hours ago

How plants become bushy, or not: New study sheds light on hormone that controls branching

11 hours ago

Elephants on the move: Mapping connections across African landscapes

11 hours ago

Study finds seasonal shifts in moral values

13 hours ago

Load comments (0)

New search engine ranks tables by title, document content, text reference

Modern aircraft emit less carbon than older aircraft, but their contrails may do more environmental harm

Scientists equip Australian sea lions with cameras to explore previously unmapped ocean habitats

Fishing disrupts squaretail grouper mating behavior, study finds

Domestication causes smaller brain size in dogs than in the wolf: Study challenges notion

Tundra vegetation to grow taller, greener through 2100, study finds

Living with a killer: How an unlikely mantis shrimp-clam association violates a biological principle

Bouncing helps people move in sync during dance, study shows

How plants become bushy, or not: New study sheds light on hormone that controls branching

Elephants on the move: Mapping connections across African landscapes

Study finds seasonal shifts in moral values

Relevant PhysicsForums posts

Creating a minimal Windows 11 Bootable USB stick for my ROG Computer

Python Socket library to create a server and client scripts

Safe, free and unlimited xls to xlsx converter?

Help solving a geometrical matching issue with Graph Neural Networks

5 GHz PC WiFi connection Cybersecurity question

Help with some optimization code for Block Matrices

New calculation approach allows more accurate predictions of how atoms ionize when impacted by high-energy electrons

New drone imagery reveals 97% of coral dead at a Lizard Island reef after last summer's mass bleaching

Why sexual violence against men by women needs to be 'called out' too

Unlocking the world of bacteria—researchers introduce new approach to make bacteria amenable to genetic engineering

New tool enables faster, more cost-effective genome editing of traits to improve agriculture sustainability

Lichen partnerships challenged by changes in the Northwoods

Hyphens in paper titles harm citation counts and journal impact factors

A big step toward the practical application of 3-D holography with high-performance computers

Combining multiple CCTV images could help catch suspects

Applying deep learning to motion capture with DeepLabCut

Training artificial intelligence with artificial X-rays

New model for large-scale 3-D facial recognition

Medical Xpress

Tech Xplore

Science X

New search engine ranks tables by title, document content, text reference

Modern aircraft emit less carbon than older aircraft, but their contrails may do more environmental harm

Scientists equip Australian sea lions with cameras to explore previously unmapped ocean habitats

Fishing disrupts squaretail grouper mating behavior, study finds

Domestication causes smaller brain size in dogs than in the wolf: Study challenges notion

Tundra vegetation to grow taller, greener through 2100, study finds

Living with a killer: How an unlikely mantis shrimp-clam association violates a biological principle

Bouncing helps people move in sync during dance, study shows

How plants become bushy, or not: New study sheds light on hormone that controls branching

Elephants on the move: Mapping connections across African landscapes

Study finds seasonal shifts in moral values

Relevant PhysicsForums posts

Related Stories

New calculation approach allows more accurate predictions of how atoms ionize when impacted by high-energy electrons

New drone imagery reveals 97% of coral dead at a Lizard Island reef after last summer's mass bleaching

Why sexual violence against men by women needs to be 'called out' too

Unlocking the world of bacteria—researchers introduce new approach to make bacteria amenable to genetic engineering

New tool enables faster, more cost-effective genome editing of traits to improve agriculture sustainability

Lichen partnerships challenged by changes in the Northwoods

Recommended for you

Hyphens in paper titles harm citation counts and journal impact factors

A big step toward the practical application of 3-D holography with high-performance computers

Combining multiple CCTV images could help catch suspects

Applying deep learning to motion capture with DeepLabCut

Training artificial intelligence with artificial X-rays

New model for large-scale 3-D facial recognition

Newsletter sign up

Donate and enjoy an ad-free experience