Skip to main navigation Skip to search Skip to main content

Universal web pages content parser

  • Silesian University of Technology
  • Institute of Theoretical and Applied Informatics of the Polish Academy of Sciences

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

5 Citations (Scopus)

Abstract

This article describes the universal web pages content parser - cross-platform application enhancing the process of data extraction from the web pages. In this implementation user friendly interface, possibility of significant automation and reusability of already created patterns had been the key elements. Moreover, the original approach to the issue of parsing the not well-formed HTML, stating the application's core, is precisely presented. Universal web pages content parser shows that the simplified web scrapping utility may be available to masses and not well-formed HTML sources may feed useful tree-like data structures as well as the well-formed ones.

Original languageEnglish
Title of host publicationComputer Networks - 19th International Conference, CN 2012, Proceedings
Pages130-138
Number of pages9
DOIs
Publication statusPublished - 2012
Event19th International Conference on Computer Networks, CN 2012 - Szczyrk, Poland
Duration: 19 Jun 201223 Jun 2012

Publication series

NameCommunications in Computer and Information Science
Volume291 CCIS
ISSN (Print)1865-0929

Conference

Conference19th International Conference on Computer Networks, CN 2012
Country/TerritoryPoland
CitySzczyrk
Period19/06/1223/06/12

Keywords

  • HTML
  • XML
  • content parser
  • web

ASJC Scopus subject areas

  • General Computer Science
  • General Mathematics

Fingerprint

Dive into the research topics of 'Universal web pages content parser'. Together they form a unique fingerprint.

Cite this