Przeskocz do nawigacji głównej Przeskocz do wyszukiwania Przeskocz do głównej treści

Universal web pages content parser

  • Silesian University of Technology
  • Institute of Theoretical and Applied Informatics of the Polish Academy of Sciences

Wyniki badań: Rozdział w książce/raport/materiał konferencyjnyWkład w konferencjęrecenzja

5 Cytowania z bazy Scopus

Abstrakt

This article describes the universal web pages content parser - cross-platform application enhancing the process of data extraction from the web pages. In this implementation user friendly interface, possibility of significant automation and reusability of already created patterns had been the key elements. Moreover, the original approach to the issue of parsing the not well-formed HTML, stating the application's core, is precisely presented. Universal web pages content parser shows that the simplified web scrapping utility may be available to masses and not well-formed HTML sources may feed useful tree-like data structures as well as the well-formed ones.

Język oryginałuangielski
Tytuł publikacji goszczącejComputer Networks - 19th International Conference, CN 2012, Proceedings
Strony130-138
Liczba stron9
Identyfikatory DOI
Status publikacjiOpublikowano - 2012
Wydarzenie19th International Conference on Computer Networks, CN 2012 - Szczyrk, Polska
Czas trwania: 19 cze 201223 cze 2012

Seria publikacji

NazwaCommunications in Computer and Information Science
Tom291 CCIS
ISSN (drukowany)1865-0929

Konferencja

Konferencja19th International Conference on Computer Networks, CN 2012
Kraj/TerytoriumPolska
MiejscowośćSzczyrk
Okres19/06/1223/06/12

Obszary tematyczne ASJC Scopus

  • Informatyka ogólna
  • Matematyka ogólna

Fingerprint

Zanurz się w tematy badawcze publikacji „Universal web pages content parser”. Razem tworzą niepowtarzalny odcisk palca.

Cytowanie