Data Scientist & Engineer
Indexed description
Practice Group / Department:
Legal Innovation, Design & Technology - LondonJob Description
Norton Rose Fulbright is a global law firm with more than 3,000 lawyers advising clients across locations in the United States, Europe, Canada, Latin America, Asia, Australia, Africa and the Middle East. We provide a full scope of legal services to the world’s preeminent corporations and financial institutions.
Our vision is to be a world class business, profitable, ambitious, cooperative and considerate, supporting our clients and people through our global business principles of Quality, Unity and Integrity.
With over 7,000 employees worldwide, our culture is the thread that connects us. Our strategy and culture are closed connected – defined by shared ambition, global collaboration and a one-team mindset. We believe pioneering work happens when people are empowered to think beyond boundaries, explore new opportunities and grow through diverse experiences. Alongside the right skills and experience, we are looking for people who are innovative, commercially minded, and motivated by the impact of the work they do – ready to share in our ambition and help shape what comes next.
Because while individuals can do well, together we achieve something extraordinary.
Role Purpose
We are building a new R&D capability focused on developing data-driven and AI-enabled products for legal services and the wider business of law.
The Data Scientist & Engineer will work within R&D alongside the Data Programme to build the foundational data capabilities required for those products: trusted, reusable data assets, reliable pipelines, data quality, lakehouse models and analytical products. These capabilities will support R&D products, AI applications, client intelligence and firm-wide decision support.
This is a hands-on hybrid data-engineering and applied-data-science role, with particular responsibility for building data products in Microsoft Fabric and Databricks. You will work with Data Programme colleagues to turn priority sources into governed, use-case-ready data assets using agreed patterns, without duplicating enterprise platform responsibilities.
You will join a small, hands-on multidisciplinary team working collaboratively across discovery, prototyping, engineering, productionisation and continuous improvement.
Key Responsibilities
-
Partner with the Data Programme, data owners and system teams to assess, access and combine priority internal and external data sources, understanding quality, permissions and fitness for use.
-
Design, build and operate reliable pipelines in Microsoft Fabric and Databricks to ingest, clean, standardise, enrich, version and publish structured, semi-structured and unstructured data as reusable data assets.
-
Create scalable lakehouse data models, schemas, data contracts and reusable features for entities such as clients, organisations, people, matters, sectors, jurisdictions, opportunities and legal topics.
-
Implement data-quality checks, lineage, provenance, observability, monitoring and change management within Fabric and Databricks workflows so data assets remain trusted over time.
-
Conduct exploratory data analysis to surface patterns, gaps, anomalies and data-quality issues, and identify opportunities for analysis, modelling or product use.
-
Develop, compare and validate appropriate statistical and machine-learning models for classification, ranking, recommendation, forecasting, anomaly detection, similarity or extraction problems.
-
Use NLP, embeddings and LLM-assisted techniques responsibly for classification, entity resolution, information extraction, clustering and relationship discovery across document-heavy data.
-
Publish curated datasets, features and analytical outputs through approved Fabric and Databricks data products, tables, APIs, search indexes, batch pipelines or product features.
-
Maintain reproducible code, tests, model and data documentation, evaluation evidence and clear statements of limitations and appropriate use.
-
Use AI agents and coding assistants responsibly to accelerate exploratory work, pipeline development, testing and documentation, while retaining ownership of analytical judgement and verification.
Initial Focus
-
Establish foundational, governed and permission-aware data assets in Microsoft Fabric and Databricks, including curated datasets, metadata, lakehouse data models and quality measures.
-
Develop reusable pipelines in Microsoft Fabric and Databricks for entity resolution, enrichment, metadata, relationship mapping and data-quality monitoring.
-
Run EDA and model development that increases the value of priority data assets for client intelligence, opportunity identification, pricing, matter analysis and other high-value use cases.
What Success Looks Like
-
The Data Programme and R&D teams can use documented, quality-controlled data assets in Fabric and Databricks without repeatedly cleaning and reconciling source data.
-
Analytical work identifies useful signals and viable opportunities, with transparent baselines, validation and limitations.
Create a free Caio profile to unlock more results and save your role and location preferences.
Unlock free search