Abstract
ETL processes are sometimes interrupted by occurrence of failure. After that, one of the interrupted extraction resumption algorithms is usually used. In this paper we present a modified Design-Resume algorithm (Redo group) enriched by the possibility of handling ETL processes containing many loading nodes. The key feature of this algorithm is that it does not impose additional overhead on normal ETL process. Extending the resumption time in some situations is the drawback of the presented Design-Resume resumption. To solve this problem we present an estimation method predicting the resumption time which is the modification of the method implemented in our previous DR/JB extraction environment. This let us decide what action should be taken: resumption or running the extraction from the beginning. The accuracy of the estimation is examined in the paper, then the estimation algorithm's drawbacks are discussed and factors making estimation imprecise are enumerated. Finally, we summarize obtained results and show our future plans.
| Original language | English |
|---|---|
| Pages | 284-292 |
| Number of pages | 9 |
| Publication status | Published - 2004 |
| Event | 15th International Conference on Systems Science - Wroclaw, Poland Duration: 7 Sept 2004 → 10 Sept 2004 |
Conference
| Conference | 15th International Conference on Systems Science |
|---|---|
| Country/Territory | Poland |
| City | Wroclaw |
| Period | 7/09/04 → 10/09/04 |
Keywords
- Data warehouse
- ETL
- Estimation
- Resumption
ASJC Scopus subject areas
- General Engineering
Fingerprint
Dive into the research topics of 'Estimation of profitability of Design-Resume/JavaBeans resumption algorithm'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver