BlueFlag LLC Logo

BlueFlag LLC

Senior Data Scientist

Posted Yesterday
In-Office or Remote
Hiring Remotely in Austin, TX, USA
Senior level
In-Office or Remote
Hiring Remotely in Austin, TX, USA
Senior level
Build and deploy production machine learning and AI solutions on Databricks and Azure. Responsibilities include data discovery, feature engineering, model development, experiment tracking, model serving, monitoring, retraining, LLM and RAG use cases, responsible AI documentation, and modernization of legacy R, Stata, and SAS workloads. The role also involves client consultation, cloud adoption, training, mentoring, and communicating results to technical and executive audiences.
The summary above was generated by AI

BlueFlag is looking for a Senior Data Scientist to support a premier data and AI platform for the Department of Veterans Affairs. You'll work directly with internal clients to turn operational problems into production machine learning and AI solutions on Databricks in Azure, and help their teams adopt modern cloud tooling. This role is hands-on. We want someone who has taken models from a notebook to production and kept them running.

What You'll Do

  • Consult with internal clients to frame business problems as analytical or ML problems, define success metrics, and set realistic scope
  • Build data products and workflows that support critical operations, from source data discovery through production deployment
  • Migrate and modernize client workloads from legacy environments (including R, Stata, and SAS) to Python and Databricks, with measurable gains in runtime, cost, or reliability
  • Develop, train, and validate ML models (classification, regression, forecasting, clustering, anomaly detection) for complex business problems
  • Engineer features and build reusable, governed feature pipelines on large datasets with PySpark and SQL
  • Track experiments, register models, and manage model versions and promotion through MLflow
  • Deploy models for batch scoring and real-time serving, and automate retraining with scheduled jobs and CI/CD pipelines
  • Monitor models in production for performance, data drift, and data quality, and set thresholds and alerts that trigger review or retraining
  • Lead AI adoption with clients, including LLM and agentic use cases such as retrieval-augmented generation (RAG), document summarization, and classification
  • Evaluate LLM and agent outputs for accuracy, groundedness, and safety before and after release
  • Document models (purpose, data, assumptions, limitations, evaluation results) so they hold up to governance and responsible AI review
  • Build reference use cases with tutorials, reference code, and training to drive adoption of cloud tooling for concrete business results
  • Host office hours and pair with client analysts and data scientists to upskill them
  • Present findings and model results clearly to technical and non-technical audiences, including leadership

Why Join BlueFlag

At BlueFlag, we're passionate about leveraging cutting-edge technology to make a real difference. You'll be at the forefront of cloud innovation, working on projects that directly impact people's lives. We offer a high-growth, entrepreneurial environment that values fresh ideas and authentic teamwork.

If you're ready to take your data engineering career to new heights and contribute to meaningful projects that push the boundaries of technology, we want to hear from you. Join BlueFlag and be part of a team that's shaping the future of AI solutions!


Requirements
  • Bachelor's degree in Engineering, Computer Science, Statistics, Mathematics, Systems, Business, or a related scientific or technical discipline, and 15+ years of experience (or commensurate experience)
  • Proficient in Python (pandas, NumPy, SciPy, scikit-learn) and advanced SQL (window functions, CTEs, query tuning) for data analysis
  • 2+ years of hands-on work on a leading cloud data platform such as Databricks, Azure, AWS, or GCP
  • Experience across the end-to-end data science workflow, from finding and assessing datasets to production deployment, including experiment tracking and model management with MLflow or similar
  • Solid grounding in traditional machine learning: supervised and unsupervised methods, gradient-boosted trees (XGBoost, LightGBM), model selection, cross-validation, and hyperparameter tuning
  • Strong applied statistics: hypothesis testing, regression, sampling, and experimental design
  • Experience working with large datasets in a distributed environment (Spark/PySpark)
  • Sound model evaluation practice: picking the right metrics, handling class imbalance, avoiding leakage, and explaining model behavior (for example SHAP or feature importance)
  • Working knowledge of large language models (LLMs) and agentic AI workflows, including prompt design and RAG patterns
  • Version control with Git and collaborative development practices (code review, branching, testing)
  • Ability to explain technical work to non-technical stakeholders and turn ambiguous requests into defined deliverables
  • US Citizen: Must be a citizen of the United States
  • Security Clearance: Must be able to obtain a public trust clearance.  Must be eligible to work in the United States.

Desired

  • 5+ years as a data scientist, shipping multiple products that run in operation
  • Working experience with Databricks in Azure, including Unity Catalog, Delta Lake, Databricks Jobs, and Databricks notebooks/Repos
  • MLOps experience across the lifecycle, such as:
    • Model Serving endpoints for real-time inference
    • Feature engineering and feature tables governed in Unity Catalog
    • CI/CD for ML with Azure DevOps or GitHub Actions, and Databricks Asset Bundles
    • Production monitoring for drift and model quality (for example Lakehouse Monitoring)
    • Champion/challenger or A/B testing of models in production
    • Automated retraining and model lineage and auditability
  • Experience with Azure AI services (Azure OpenAI, Azure Machine Learning) or Mosaic AI (Vector Search, Agent Framework)
  • Prior experience shipping products that use LLMs or AI agents, including evaluation and guardrails
  • 2+ years building visual insights (dashboards and reports) that support operational needs, with Power BI, Databricks AI/BI dashboards, or Tableau
  • Prior experience with R, Stata, or SAS
  • Experience refactoring R, Stata, or SAS codebases to Python, including validating that results match
  • Experience with VA or federal healthcare data, and handling PHI/PII under federal privacy and security requirements
  • Familiarity with federal AI governance and responsible AI practices (bias testing, model documentation, human oversight)
  • Experience training or mentoring analysts and data scientists
  • Master's or PhD in a quantitative field

Benefits
  • Competitive salary
  • Generous annual leave and paid holidays
  • Comprehensive group health and dental plans
  • 401(k) with company match
  • Life insurance and AD&D coverage
  • Ongoing training and professional development opportunities

Similar Jobs

13 Days Ago
Remote
United States
170K-200K Annually
Senior level
170K-200K Annually
Senior level
Aerospace • Artificial Intelligence • Computer Vision • Software • Analytics • Defense • Big Data Analytics
Leads development of statistical and machine learning models for time-series, geospatial, kinematic, and radar data. Responsibilities include feature engineering, forecasting, anomaly detection, data integration, ETL, visualization, algorithm evaluation, simulation, dataset creation, and stakeholder communication. The role also contributes to mission-focused software applications, APIs, data pipelines, documentation, code reviews, and agile development while requiring an active Top Secret or CBP/DHS suitability clearance.
Top Skills: Agile ScrumGitHdfsKanbanKerasMatplotlibNumpyPandasPostgresPythonPyTorchScikit-LearnScipyTensorFlow
21 Days Ago
Remote
United States
170K-200K Annually
Senior level
170K-200K Annually
Senior level
Aerospace • Artificial Intelligence • Computer Vision • Software • Analytics • Defense • Big Data Analytics
Lead development of statistical and machine learning models for time-series, geospatial, kinematic, and radar data. Build analytic solutions, training datasets, ETL workflows, forecasting and anomaly-detection models, simulations, visualizations, and performance metrics. Integrate multi-source sensor data, enhance algorithms, document processes, and collaborate with engineers, analysts, mission partners, and stakeholders in an agile environment.
Top Skills: Agile ScrumGitHdfsKanbanKerasMatplotlibNumpyPandasPostgresPythonPyTorchScikit-LearnScipyTensorFlow
6 Hours Ago
Remote
United States
131K-188K Annually
Senior level
131K-188K Annually
Senior level
Greentech • Energy
Own and deliver production machine learning workstreams for utility bill extraction, forecasting, and audit/anomaly detection. Build and evaluate models, establish technical approaches for ambiguous problems, write production-grade Python, and partner with engineering on integration and monitoring. The role also involves applying LLMs and agentic workflows, automating team processes, prioritizing investments, and communicating model behavior and recommendations to technical and nontechnical stakeholders.
Top Skills: Agentic WorkflowsLarge Language ModelsMachine LearningNumpyPandasPythonScikit-LearnSQL

What you need to know about the Austin Tech Scene

Austin has a diverse and thriving tech ecosystem thanks to home-grown companies like Dell and major campuses for IBM, AMD and Apple. The state’s flagship university, the University of Texas at Austin, is known for its engineering school, and the city is known for its annual South by Southwest tech and media conference. Austin’s tech scene spans many verticals, but it’s particularly known for hardware, including semiconductors, as well as AI, biotechnology and cloud computing. And its food and music scene, low taxes and favorable climate has made the city a destination for tech workers from across the country.

Key Facts About Austin Tech

  • Number of Tech Workers: 180,500; 13.7% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Dell, IBM, AMD, Apple, Alphabet
  • Key Industries: Artificial intelligence, hardware, cloud computing, software, healthtech
  • Funding Landscape: $4.5 billion in VC funding in 2024 (Pitchbook)
  • Notable Investors: Live Oak Ventures, Austin Ventures, Hinge Capital, Gigafund, KdT Ventures, Next Coast Ventures, Silverton Partners
  • Research Centers and Universities: University of Texas, Southwestern University, Texas State University, Center for Complex Quantum Systems, Oden Institute for Computational Engineering and Sciences, Texas Advanced Computing Center

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account