An Effective Way of Storing and Accessing Very Large Transition Matrices Using Multi-core CPU and GPU Architectures