Because of the problem of not able to recognize the start or end of any record in JSON, it makes it difficult to parse it, this brings me to work on a project of tokenizing the JSON tokens for large data File using Hadoop MapReduce.
The Open Source Project called ElephantBird which contains utilities which will be helpful in working for MapReduce