Parallel Secondo Using Hadoop Starting with release 3.3, it is possible to use Secondo for parallel processing of queries. Hadoop needs to be installed together with Secondo. Queries in Secondo's executable language can be embedded into Hadoop Map or Reduce steps. Essentially Hadoop is used as a distributed operating system that assigns tasks to Secondo instances on different computers and supervises their execution. This approach, as is known for the MapReduce approach, is highly fault-tolerant and suitable for large networks of hundreds of computers.
Parallel queries can be formulated in one Secondo system as a query in executable language containing hadoopMap and hadoopReduce operations. The following documentation is available:
- Example: How to Write Parallel Queries in Parallel Secondo
- User Guide For Parallel Secondo
- J. Lu and R.H. Güting, Simple and Efficient Coupling of Hadoop With a Database Engine. Fernuniversität in Hagen, Informatik-Report 366 - 10/2012.
Last Changed: 2012-11-10