In [11]:
def split_data(input_data,  tokenizer, max_length = 2048):
    # Function to decode and split data
    def split_entry(entry, max_length):
        tokens = tokenizer.encode(entry)
        current_length = 0
        current_part_tokens = []
        new_entries = []

        for token in tokens:
            if current_length + 1 > max_length:
                current_part_text = tokenizer.decode(current_part_tokens, skip_special_tokens=True)
                new_entries.append(f'<system>\n{current_part_text}')
                current_part_tokens = [token]
                current_length = 1
            else:
                current_part_tokens.append(token)
                current_length += 1
        
        if current_part_tokens:
            current_part_text = tokenizer.decode(current_part_tokens, skip_special_tokens=True)
            new_entries.append(f'<system>\n{current_part_text}')
        
        return new_entries

    # Process each entry in the input data
    processed_data = []
    for entry in input_data:
        tokens = tokenizer.encode(entry)
        if len(tokens) > 2048:
            processed_data.extend(split_entry(entry))
        else:
            processed_data.append(f'<system>\n{entry}')
    
    return processed_data

In [2]:
from datasets import DatasetDict, load_dataset
import re
import random
from multiprocessing import cpu_count
from transformers import AutoTokenizer
import torch
import config

  from .autonotebook import tqdm as notebook_tqdm


In [5]:
tokenizer = AutoTokenizer.from_pretrained(config.model_name)
# set pad_token_id equal to the eos_token_id if not set
if tokenizer.pad_token_id is None:
  tokenizer.pad_token_id = tokenizer.eos_token_id

In [6]:
test_text = " <|system|>\n</s>\n<|user|>\nWrite a compelling mystery story set in a vineyard, where a seasoned detective investigates a murder with twists and turns that will keep the reader engaged until the very end. Add complex characters, multiple suspects, and red herrings to create suspense and challenge the detective's deductive reasoning. Use vivid descriptive language to paint a picture of the vineyard setting, its wine-making process, and the people who live and work there. Make sure to reveal clues and motives gradually, and create a satisfying resolution that ties up all loose ends.</s>\n<|assistant|>\nDetective Jameson had been called to the vineyard on the outskirts of town to investigate a murder. The sun was setting, casting long shadows over the grape vines, and the air was heavy with the sweet scent of fermented grapes. The body was lying in the middle of the vineyard, surrounded by broken grape vines and a scattering of grapes. The victim's throat had been slashed, and there were bruises around the neck.\n\nAs Jameson began his investigation, he met the owner of the vineyard, a tall, thin man with a stern face. The owner didn't seem overly concerned about the murder on his property and was reluctant to answer any questions that Jameson had. He was also dismissive of the other workers, including the winemaker, who appeared visibly distraught.\n\nJameson soon discovered that the victim, a young woman named Sophia, had been working at the vineyard for the past year. Sophia was well-liked by the other workers, who had nothing but positive things to say about her. She was hardworking and always had a smile on her face. As Jameson dug deeper, he found that Sophia had possibly been having an affair with the owner, who was married. The owner denied any involvement, but Jameson was suspicious, especially when he found a note in Sophia's pocket that hinted at a secret relationship.\n\nFurther investigation uncovered more potential suspects—a worker who had been fired earlier that day, a jealous ex-boyfriend, and even a rival vineyard owner. With so many suspects, Jameson had to be careful, and he wasn't sure who he could trust. The winemaker, who had initially seemed distraught, also didn't seem to be telling Jameson everything he knew.\n\nAs Jameson pieced together the clues, he found that the murder had been premeditated. The killer had used the grape vines to hide their tracks, and the bruises on Sophia's neck indicated she had been strangled before her throat was slashed. Jameson knew he was getting closer to solving the case, and he sensed the killer was getting nervous.\n\nThe final twist came when the winemaker confessed to the murder. He had been in love with Sophia and had become jealous when he discovered she had been seeing the owner. He had planned the murder carefully, knowing that Sophia had been working alone that day. Jameson arrested the winemaker, relieved that he had solved the case and that justice would be served.\n\nAs the sun rose over the vineyard the next day, the workers returned to their routine of harvesting grapes, but Jameson knew that the tranquility of the vineyard had been shattered by a brutal murder.</s>\n<|user|>\nWow, that was a great mystery story! I really enjoyed the twists and turns, and the characters were interesting. Can you give me more detailed descriptions of the vineyard setting and the wine-making process? I want to be fully immersed in the story.</s>\n<|assistant|>\nOf course! As Detective Jameson arrived at the vineyard, he was struck by the natural beauty of the surroundings. The rolling hills were covered in rows upon rows of grapevines, their leaves rustling gently in the breeze. The vineyard was divided into different sections, each one dedicated to a specific type of grape. Jameson could see the giant barrels inside the winery, where the grapes would be crushed, fermented, and aged until they became a fine wine.\n\nThe grapes themselves were a sight to behold—some were deep burgundy, others were a bright green, and some were shades in between. As Jameson walked through the vineyard, he could smell the sweet aroma of grapes hanging heavily from the vines. He saw workers picking the grapes carefully, placing them in baskets and moving them to the winery for processing. The winery itself was a sprawling building with large metal tanks, wooden barrels, and machinery scattered throughout the space.\n\nJameson had always been fascinated by the art of winemaking. He knew it was a delicate process that required patience, knowledge, and skill. It was a precise balance of science and art—the grapes had to be harvested at the right time, crushed in the right way, and aged properly to produce the perfect wine. As he learned more about the wine-making process, Jameson became convinced that the careful attention to detail could uncover critical clues to solve the murder mystery.\n\nThroughout his investigation, Jameson noticed other aspects of the vineyard that added to its charm—the sprawling estate with its mustard yellow stucco walls, the vintage bicycles parked near the winery, the creaky wooden tasting room, and the buzzing of bees that drifted through the vineyards.\n\nAll these sights and smells contributed to the peaceful setting of the vineyard, making it even more shocking when Jameson stumbled upon a gruesome murder. The picturesque vineyard was now a crime scene, and Jameson knew that he had his work cut out for him to catch the killer.</s>\n<|user|>\nThat was an amazing description! It really transported me into the vineyard with Detective Jameson. Can you tell me more about the different types of grapes that were grown in the vineyard? I'm curious to know which ones were used to create the murder victim's wine.</s>\n<|assistant|>\nYes, the vineyard grew various types of grapes, each with its unique flavor profile and characteristics. The vineyard had vines of red grapes, such as Pinot Noir, Cabernet Franc, Cabernet Sauvignon, and Merlot. These grapes added depth, structure, and complexity to the wine, and were often blended to create different red wine blends.\n\nThe vineyard also had vines of white grapes, such as Chardonnay, Sauvignon Blanc, and Viognier. These grapes were used to produce white wine, which was light, refreshing, and crisp. Viognier, in particular, was known to be used to create a small batch, highly sought-after white wine from the vineyard.\n\nSophia, the victim, was known to have been a specialist in selecting the perfect blend of grapes for each bottle of wine. She had a unique talent for detecting the subtle flavors and nuances of each grape and was often tasked with creating the vineyard's signature wines.\n\nWhen Detective Jameson examined the grape vines around the scene of the crime, he noticed that some of the vines had been uprooted and damaged. He also saw that the grapes in the area were Merlot grapes, which were used to create the vineyard's award-winning red wine.\n\nAs Jameson dove deeper into the investigation, he uncovered a plot to sabotage the Merlot grapes. Someone had deliberately damaged the vines, destroying the grapes' flavor and texture, and ultimately ruining the wine. Sophia had been investigating the sabotage, and that could have led to her murder.\n\nDetective Jameson quickly realized that the Merlot grapes were a crucial clue in the murder investigation. He interviewed every worker present at the vineyard, looked through the vineyard's production logs, and analyzed samples of each grape variety to identify potential suspects. The hunt for the killer led Jameson on a roller-coaster ride of clues and red herrings, and ultimately to the arrest of the winemaker.</s>\n"

In [18]:
a = test_text.split('</s>\n<|user|>\n')[2]

In [21]:
len(tokenizer.encode(a))

566

In [None]:
for segment in test_text.split('</s>\n<|user|>\n')[1:]:
    sequence = ''
    while(tokenizer.encode(sequence) < 2048):
        
    sequence += segment
    

In [27]:
from datasets import load_from_disk

train = load_from_disk("/gpfs-volume/SFT/train_data")

In [30]:
train['text'][0]

"<|system|>\n</s>\n<|user|>\nHow does the location of the Sydney Conservatorium of Music impact the academic and professional opportunities available to music students, and how does the conservatorium support student engagement with the music industry in Australia?</s>\n<|assistant|>\nThe location of the Sydney Conservatorium of Music, which is situated in the heart of Sydney's cultural precinct, impacts both the academic and professional opportunities available to music students. The conservatorium is located near several major performing arts venues and organizations, including the Sydney Opera House, the Australian Broadcasting Corporation, and the Sydney Symphony Orchestra, providing students with easy access to performances, rehearsals, and networking opportunities.\n\nOne of the primary ways the conservatorium supports student engagement with the music industry in Australia is through its strong industry connections. The conservatorium has established partnerships with numerous p

In [26]:
print(split_data(test_text))

["<|system|>\n</s>\n<|user|>\nWrite a compelling mystery story set in a vineyard, where a seasoned detective investigates a murder with twists and turns that will keep the reader engaged until the very end. Add complex characters, multiple suspects, and red herrings to create suspense and challenge the detective's deductive reasoning. Use vivid descriptive language to paint a picture of the vineyard setting, its wine-making process, and the people who live and work there. Make sure to reveal clues and motives gradually, and create a satisfying resolution that ties up all loose ends.</s>\n<|assistant|>\nDetective Jameson had been called to the vineyard on the outskirts of town to investigate a murder. The sun was setting, casting long shadows over the grape vines, and the air was heavy with the sweet scent of fermented grapes. The body was lying in the middle of the vineyard, surrounded by broken grape vines and a scattering of grapes. The victim's throat had been slashed, and there wer