<div>Remote | Data Scientist & Quantitative Analyst - $55-$85/hour <br/><br/><strong>We are sharing a specialised full-time consulting opportunity for experienced data scientists and quantitative analysts with strong expertise in statistical analysis, data cleaning, method comparison, reproducible research, and evidence-based reporting.</strong><br/><br/>This role supports the development of advanced agentic evaluation benchmarks for frontier AI models. Selected professionals will create realistic data-analysis challenges, develop reproducible reference notebooks, evaluate model-generated analyses, and identify where statistical reasoning, interpretation, or reporting falls short of professional standards.<br/><br/><strong>Key Responsibilities</strong><br/><br/><strong>Data Analysis Task Design</strong><br/><ul><li>Create realistic analytical tasks based on professional data science and quantitative research workflows</li><li>Develop assignments involving messy data, anomaly detection, correlation analysis, hypothesis testing, and method comparison</li><li>Design complex, multi-step problems requiring statistical judgment and careful interpretation</li><li>Ensure tasks include realistic constraints, datasets, assumptions, and decision-making objectives</li></ul><br/><strong>Reproducible Notebook Development</strong><br/><ul><li>Complete reference analyses using Jupyter Notebook or Google Colab</li><li>Build clear and reproducible workflows using Python, pandas, NumPy, and related libraries</li><li>Document data-cleaning decisions, calculations, statistical methods, and analytical conclusions</li><li>Validate intermediate results, spot checks, visualisations, and final recommendations</li></ul><br/><strong>Statistical Method Comparison</strong><br/><ul><li>Design fair comparisons between analytical models, algorithms, or statistical approaches</li><li>Evaluate performance using appropriate metrics, manual checks, and sensitivity analyses</li><li>Identify methodological trade-offs, limitations, and sources of uncertainty</li><li>Produce recommendations supported by transparent quantitative evidence</li></ul><br/><strong>AI Model Evaluation</strong><br/><ul><li>Review model-generated analyses for statistical accuracy, methodological rigour, and sound interpretation</li><li>Verify whether calculations, correlations, hypotheses, and conclusions are supported by the data</li><li>Identify coding errors, unsupported assumptions, misleading summaries, and analytical shortcuts</li><li>Explain where and why model outputs fail to meet professional data-analysis standards</li></ul><br/><strong>Research Collaboration</strong><br/><ul><li>Work closely with researchers, task authors, and fellow quantitative specialists</li><li>Compare evaluation decisions to maintain consistent benchmark standards</li><li>Refine tasks, reference notebooks, and grading criteria based on testing outcomes</li><li>Document recurring model weaknesses and opportunities for stronger evaluation coverage</li></ul><br/><strong>Ideal Profile</strong><br/><br/><strong>Strong candidates may have:</strong><br/><ul><li>At least 1 year of experience in data science, quantitative analysis, research engineering, or another research-intensive analytical role</li><li>Deep hands-on experience with data cleaning, statistical correlation, hypothesis testing, and interpretation</li><li>Strong proficiency in Python, including pandas, NumPy, or comparable analytical libraries</li><li>Experience using Jupyter Notebook or Google Colab for analysis and reporting</li><li>Working familiarity with Git and reproducible analytical workflows</li><li>Ability to communicate complex quantitative findings clearly to technical and non-technical decision-makers</li><li>Strong attention to detail and confidence working through ambiguous, open-ended problems</li><li>Reliable availability for approximately 35 hours per week</li></ul><br/><strong>Educational Background</strong><br/><ul><li>A master's degree or PhD in statistics, data science, mathematics, economics, computer science, engineering, or another quantitative discipline is highly relevant</li><li>Equivalent practical experience in a research-heavy analytical field may also be considered</li><li>Academic or professional research involving statistical modelling, experimentation, or large-scale data analysis may strengthen an application</li><li>Publications, technical reports, open-source work, or impactful analytical projects may also be valuable</li></ul><br/><strong>Nice to Have</strong><br/><ul><li>Experience in AI training, model evaluation, or benchmark development</li><li>Background authoring analytical tasks, reference solutions, or grading rubrics</li><li>Familiarity with anomaly detection, experimental design, or comparative model evaluation</li><li>Experience conducting manual spot checks and validating automated analyses</li><li>Knowledge of statistical modelling, machine learning, or scientific computing</li><li>Familiarity with agentic AI systems and multi-step model evaluations</li><li>Experience reviewing notebooks, code, or analyses prepared by other professionals</li><li>Strong ability to identify subtle statistical errors and unsupported conclusions</li></ul><br/><strong>Why This Opportunity</strong><br/><ul><li>Apply advanced data science and quantitative analysis expertise to frontier AI evaluation</li><li>Design realistic tasks grounded in professional analytical workflows</li><li>Help improve how AI systems reason through statistics, data quality, and method comparison</li><li>Work across Python, reproducible notebooks, model evaluation, and evidence-based reporting</li><li>Collaborate closely with researchers and other quantitative specialists</li><li>Participate in a structured full-time remote role with competitive hourly compensation</li></ul><br/><strong>Contract Details</strong><br/><ul><li>Full-time W-2 contingent employment opportunity</li><li>Fully remote within the United States</li><li>Expected commitment of approximately 35 hours per week</li><li>Competitive rates between $55-$85 per hour depending on expertise and project scope</li><li>Individual tasks may require one to two days of focused analysis and implementation</li><li>Work may include task design, data cleaning, statistical analysis, notebook development, AI output evaluation, and technical reporting</li><li>Engagement scope and duration may evolve according to project requirements and performance</li></ul></div>