In [1]:
"""
1. Search for an open opportunity at your possible potential clients
2. Search for relevant information from the opening that might act as a good starting point for your business with this company
3. Go through all of your company's projects and find relevant projects
4. Get important to share information about these projects
5. Write a compelling cold email about how your firm can be of great service to your potential client
6. Repeat for ALL potential clients
"""

"\n1. Search for an open opportunity at your possible potential clients\n2. Search for relevant information from the opening that might act as a good starting point for your business with this company\n3. Go through all of your company's projects and find relevant projects\n4. Get important to share information about these projects\n5. Write a compelling cold email about how your firm can be of great service to your potential client\n6. Repeat for ALL potential clients\n"

In [None]:
# Web Scraping

# Relevant information

# Chromadb query

# Email Generation

In [2]:
from langchain_community.document_loaders import WebBaseLoader

loader = WebBaseLoader("https://meinestadt.jobs.personio.de/job/1757093?language=en&display=de")
page_data = loader.load().pop().page_content
#print(page_data)

USER_AGENT environment variable not set, consider setting it to identify your requests.


In [3]:
from langchain_groq import ChatGroq

llm = ChatGroq(
    model="DeepSeek-R1-Distill-Llama-70b",
    temperature=0,
    max_tokens=None,
    timeout=None,
    max_retries=2,
    # other params...
)

#response = llm.invoke("Who is the best footballer of all times?, no preamble")
#print(response.content)

In [4]:
from langchain_core.prompts import PromptTemplate

prompt_extract = PromptTemplate.from_template(
    """
    I will give you scraped text from the job posting. 
    Your job is to extract the job details & requirements in a JSON format containing the following keys: 'role', 'experience', 'skills', and 'description'. 
    Only return valid JSON. No preamble, please.
    Here is the scraped text: {page_data}
    """
)

chain_extract = prompt_extract | llm 
response = chain_extract.invoke(input={'page_data':page_data})
print(type(response.content))
print(response.content)

<class 'str'>
<think>
Okay, I need to extract job details from the given text and structure them into a JSON with specific keys: role, experience, skills, and description. Let me go through the text step by step.

First, the role is clearly mentioned as "Team Lead Software Engineering (m/f/x)" or "Team Lead Software Engineering (m/w/d)". I'll use the first occurrence.

Next, for experience, the text says "Mehrjährige Erfahrung in der Softwareentwicklung, davon mehr als 2 Jahre in Führungspositionen" which translates to several years of experience with over 2 years in leadership. I'll note that as "Several years of software engineering experience, including more than 2 years in leadership roles."

Now, skills. The text lists a few points. I'll translate and list them: solid technical expertise in designing complex software solutions, preferably in an app environment; approaching problems with a focus on simplicity; passion for people and leadership without micromanagement; building rela

In [5]:
from langchain_core.output_parsers import JsonOutputParser

json_parser = JsonOutputParser()
json_response = json_parser.parse(response.content)
print(type(json_response))

<class 'dict'>


In [8]:
import csv

def read_csv_file(file_path):
    data = []
    with open(file_path, 'r') as file:
        csv_reader = csv.reader(file)
        # Skip the header row
        next(csv_reader)
        for row in csv_reader:
            # Separate technical skills (list) and project link (string)
            skills = tuple(row[:-1])  # Exclude the last element (project link)
            project_link = row[-1]  # Get the last element (project link)
            data.append((skills, project_link))  # Create a tuple with skills and link
    return data

# Example usage:
file_path = 'sample_portfolio.csv'  # Replace with your actual file path
data = read_csv_file(file_path)

for skills, project_link in data:
    print(skills, project_link)
    #print(project_link)

('Python', ' SQL', ' Pandas')  https://github.com/user/project1
('SQL', ' Python', ' Airflow')  https://github.com/user/project2
('PySpark', ' Spark SQL', ' Delta Lake')  https://github.com/user/project3
('Machine Learning', ' Deep Learning', ' TensorFlow')  https://github.com/user/project4
('Data Engineering', ' ETL', ' ELT')  https://github.com/user/project5
('Cloud Platforms (AWS', ' GCP', ' Azure)')  https://github.com/user/project6
('Data Warehousing', ' Data Modeling', ' DBT')  https://github.com/user/project7
('Data Visualization', ' Power BI', ' Tableau')  https://github.com/user/project8
('MLOps', ' MLflow', ' Kubeflow')  https://github.com/user/project9
('Natural Language Processing (NLP)', ' NLTK', ' spaCy')  https://github.com/user/project10
('Computer Vision', ' OpenCV', ' TensorFlow')  https://github.com/user/project11
('Time Series Analysis', ' Forecasting', ' Prophet')  https://github.com/user/project12
('Data Cleaning', ' Data Wrangling', ' Pandas')  https://github.com

In [9]:
# Insert data into Vector Database

import uuid
import chromadb

client = chromadb.PersistentClient('vectorstore')
collection = client.get_or_create_collection(name='portfolio_links')

if not collection.count():
    for skills, project_url in data:
        collection.add(
            documents=str(skills),
            metadatas={'portfolio_url':project_url},
            ids=[str(uuid.uuid4())]
        )

In [10]:
json_response['skills']

['Solid technical expertise in designing and developing large-scale, complex software solutions, preferably in an app environment.',
 'Approach complex issues with an open mindset and a focus on achieving simplicity.',
 'Passionate about people, offering support and guidance to software engineers while avoiding micromanagement.',
 'Skilled at rapidly establishing relationships with stakeholders and maintaining a solution-focused approach.',
 'Capable of building high-performing, self-sufficient teams that effectively balance speed, quality, and risk.']

In [11]:
portfolio_urls = collection.query(query_texts=json_response['skills'], n_results=2)
portfolio_urls

{'ids': [['ed25f5bf-756b-41c5-9289-973d7577221e',
   '606a09e0-5804-46ec-868e-c71b5e3de595'],
  ['7e193151-a888-407e-a23b-1e6b7a7bed4a',
   'ed25f5bf-756b-41c5-9289-973d7577221e'],
  ['ed25f5bf-756b-41c5-9289-973d7577221e',
   '27bcbae2-0fe1-4f95-ac00-0470c9841ad4'],
  ['e049500f-f635-4ffd-a4f8-cd42e087cdbe',
   'ed25f5bf-756b-41c5-9289-973d7577221e'],
  ['e049500f-f635-4ffd-a4f8-cd42e087cdbe',
   '27bcbae2-0fe1-4f95-ac00-0470c9841ad4']],
 'embeddings': None,
 'documents': [["('Feature Engineering', ' Scikit-learn')",
   "('Data Warehousing', ' Data Modeling', ' DBT')"],
  ["('Data Ethics', ' Privacy', ' GDPR')",
   "('Feature Engineering', ' Scikit-learn')"],
  ["('Feature Engineering', ' Scikit-learn')",
   "('Data Engineering', ' ETL', ' ELT')"],
  ["('Data Governance', ' Data Quality', ' Metadata Management')",
   "('Feature Engineering', ' Scikit-learn')"],
  ["('Data Governance', ' Data Quality', ' Metadata Management')",
   "('Data Engineering', ' ETL', ' ELT')"]],
 'uris': None

In [12]:
job_description = json_response['description']

In [13]:
email_prompt = PromptTemplate.from_template(
    """
    I will give you a role and a task that you have to perform in that specific role.
    Your Role: Your name is Hassan, You are an incredible business development officer who knows how to get clients. You work for X Consulting firm, your firm works with all sorts of IT clients and provide solutions in the domain of Data Science and AI. 
    X AI focuses on efficient tailored solutions for all clients keeping costs down. 
    Your Job: Your Job is to write cold emails to clients regarding the Job openings that they have advertised. Try to pitch your clients with an email hook that opens a conversation about a possibility of working with them. Add the most relevant portfolio URLs from
    the following (shared below) to showcase that we have the right expertise to get the job done. 
    I will now provide you with the Job description and the portfolio URLs:
    JOB DESCRIPTION: {job_description}
    ------
    PORTFOLIO URLS: {portfolio_urls}
    """
)

chain_email = email_prompt | llm
response = chain_email.invoke({'job_description':job_description, 'portfolio_urls':portfolio_urls})
print(response.content)

<think>
Okay, so I need to help Hassan write a cold email to a potential client. The client has a job opening for a Software Development Team Lead, and Hassan's firm, X Consulting, wants to pitch their services. 

First, I should understand the job description. The role involves people leadership, delivery quality, and management. The person will lead a software development team, ensure high-quality code, collaborate with the Product team, and oversee feature development. The company values forward-thinking, teamwork, determination, candor, and dedication.

Now, looking at the portfolio URLs provided. There are several projects listed, each with different technologies and areas of expertise. For example, some projects focus on Feature Engineering with Scikit-learn, Data Warehousing, Data Modeling with DBT, Data Ethics, Privacy, GDPR, Data Governance, Data Quality, Metadata Management, Data Engineering with ETL and ELT.

I need to pick the most relevant portfolio URLs that align with th