H-RADIC: A Fault Tolerance Framework for Virtual Clusters on Multi-Cloud Environments

Even though the cloud platform promises to be reliable, several availability incidents prove that it is not. How can we be sure that a parallel application finishes it´s execution even if a site is affected by a failure? This paper presents H-RADIC, an approach based on RADIC architecture, that exec...

Descripción completa

Detalles Bibliográficos
Autores principales: Royo, Ambrosio, Villamayor, Jorge, Castro-León, Marcela, Rexachs del Rosario, Dolores, Luque Fadón, Emilio
Formato: Articulo
Lenguaje:Inglés
Publicado: 2018
Materias:
Acceso en línea:http://sedici.unlp.edu.ar/handle/10915/71655
http://journal.info.unlp.edu.ar/JCST/article/view/1150/909
Aporte de:
id I19-R120-10915-71655
record_format dspace
institution Universidad Nacional de La Plata
institution_str I-19
repository_str R-120
collection SEDICI (UNLP)
language Inglés
topic Ciencias Informáticas
Fault-tolerance
nube
tolerancia a fallos
computación de altas prestaciones
cloud
high- performance computing
RADIC
spellingShingle Ciencias Informáticas
Fault-tolerance
nube
tolerancia a fallos
computación de altas prestaciones
cloud
high- performance computing
RADIC
Royo, Ambrosio
Villamayor, Jorge
Castro-León, Marcela
Rexachs del Rosario, Dolores
Luque Fadón, Emilio
H-RADIC: A Fault Tolerance Framework for Virtual Clusters on Multi-Cloud Environments
topic_facet Ciencias Informáticas
Fault-tolerance
nube
tolerancia a fallos
computación de altas prestaciones
cloud
high- performance computing
RADIC
description Even though the cloud platform promises to be reliable, several availability incidents prove that it is not. How can we be sure that a parallel application finishes it´s execution even if a site is affected by a failure? This paper presents H-RADIC, an approach based on RADIC architecture, that executes parallel applications protected by RADIC in at least 3 different virtual clusters or sites. The execution state of each site is saved periodically in another site and it is recovered in case of failure. The paper details the configuration of the architecture and the experiment´s results using 3 clusters running NAS parallel applications protected with DMTCP, a very well-known distributed multi-threaded checkpoint tool. Our experiments show that by adding a cluster protector it will be possible to implement the next level in the hierarchy, where the first level in the RADIC hierarchy works as an observer at a site level. In adition, the experiments showed that the protection implementation is out of the critical path of the application and it depends on the utilized resources.
format Articulo
Articulo
author Royo, Ambrosio
Villamayor, Jorge
Castro-León, Marcela
Rexachs del Rosario, Dolores
Luque Fadón, Emilio
author_facet Royo, Ambrosio
Villamayor, Jorge
Castro-León, Marcela
Rexachs del Rosario, Dolores
Luque Fadón, Emilio
author_sort Royo, Ambrosio
title H-RADIC: A Fault Tolerance Framework for Virtual Clusters on Multi-Cloud Environments
title_short H-RADIC: A Fault Tolerance Framework for Virtual Clusters on Multi-Cloud Environments
title_full H-RADIC: A Fault Tolerance Framework for Virtual Clusters on Multi-Cloud Environments
title_fullStr H-RADIC: A Fault Tolerance Framework for Virtual Clusters on Multi-Cloud Environments
title_full_unstemmed H-RADIC: A Fault Tolerance Framework for Virtual Clusters on Multi-Cloud Environments
title_sort h-radic: a fault tolerance framework for virtual clusters on multi-cloud environments
publishDate 2018
url http://sedici.unlp.edu.ar/handle/10915/71655
http://journal.info.unlp.edu.ar/JCST/article/view/1150/909
work_keys_str_mv AT royoambrosio hradicafaulttoleranceframeworkforvirtualclustersonmulticloudenvironments
AT villamayorjorge hradicafaulttoleranceframeworkforvirtualclustersonmulticloudenvironments
AT castroleonmarcela hradicafaulttoleranceframeworkforvirtualclustersonmulticloudenvironments
AT rexachsdelrosariodolores hradicafaulttoleranceframeworkforvirtualclustersonmulticloudenvironments
AT luquefadonemilio hradicafaulttoleranceframeworkforvirtualclustersonmulticloudenvironments
AT royoambrosio hradicunasoluciondetoleranciaafallosparaclusteresvirtualesenambientesmultinube
AT villamayorjorge hradicunasoluciondetoleranciaafallosparaclusteresvirtualesenambientesmultinube
AT castroleonmarcela hradicunasoluciondetoleranciaafallosparaclusteresvirtualesenambientesmultinube
AT rexachsdelrosariodolores hradicunasoluciondetoleranciaafallosparaclusteresvirtualesenambientesmultinube
AT luquefadonemilio hradicunasoluciondetoleranciaafallosparaclusteresvirtualesenambientesmultinube
bdutipo_str Repositorios
_version_ 1764820482682519553