Vertex Macro | Trader Hub · Analysis report · July 2026
Vertex Macro | More Factors, Less Alpha
More Factors, Less Alpha
Cloud computing and generative AI make it easier for financial institutions to discover candidate signals, and also make it easier for them to produce statistical illusions at high speed
Note: This article discusses quantitative research, technical architecture, and risk governance. It does not constitute a securities recommendation, investment advice, or a return commitment.
Quantitative investing once suffered from a shortage of data. Today, it is more likely to be drowned by data.
Market prices, financial statements, regulatory filings, news, search records, social media, management remarks, and ESG materials continuously generate computable variables. Cloud platforms make storing and processing these materials cheaper and cheaper, and generative AI rapidly turns text that was once difficult to quantify into labels, scores, and vectors.
The result looks like a research boom. Genuine Alpha has not increased at the same speed.
Financial institutions can now test hundreds of candidate factors within a few days, and can also let machines extract sentiment, risk, and strategic change from thousands of reports. But discovering a variable that is correlated with historical returns, and discovering a trading edge that can be allocated capital, are two different things.
A variable warehouse can be filled with ideas. A factor factory must reject most of them.
Its work is not to produce the largest number of signals, but to use economic explanation, out-of-sample testing, transaction costs, and risk governance to identify the few signals that deserve to enter a live portfolio. The scarcest capability of a modern quantitative institution is therefore not computation, but the disciplined negation of computational results.
A number is not yet a factor
Financial research likes to give variables names. The price-to-book ratio is called a value factor, past returns become a momentum factor, the debt-to-equity ratio represents leverage, and the tone of the news becomes a sentiment factor.
Naming does not create investment value.
A variable qualifies as a candidate factor only after it answers several questions: why might it affect future returns? Who supplies this return in the trade? Is the relationship a mispricing, or compensation for risk? Why have other investors not eliminated it quickly?
The better long-term performance of cheap stocks may come from investors being overly pessimistic about distressed firms, or from the market's preference for growth narratives; it may also simply be that cheaper firms bear more operating and financing risk. The three explanations can produce similar backtests, yet they correspond to entirely different ways of losing money.
If the return comes from a behavioral bias, market learning will weaken it. If the return comes from risk compensation, a large loss in an economic downturn may not be a model failure, but the nature of the factor. If the result comes from data mining, out-of-sample testing will usually carry out the execution for the market.
The first document of a factor study should therefore not be a return curve, but an economic-hypothesis memo. It should state the mechanism, the holding period, the applicable market, the potential other side of the trade, the principal risks, capacity, and the conditions of failure.
Looking at the result first and then adding a story is the cheapest literary exercise in research, and the most expensive habit in capital allocation.
The prosperity of the zoo
Factor research has already formed a crowded zoo. Academic papers, open-source code, and standard databases have lowered the research threshold, and machine learning has further expanded the combinations of variables and parameters that can be tested.
This produces a simple and dangerous statistical fact: the more one tests, the more accidental successes one will also obtain.
If researchers test hundreds of invalid variables, some of them will always reach the usual significance threshold. Changing the sample period, the ranking method, the holding horizon, the treatment of extremes, and the industry-neutralization method will further increase the chance of finding the "best version." The t-statistic finally presented may only be the survivor among a large number of failed experiments.
A single significant result detached from the number of studies is like a lottery winner detached from the number of people who bought tickets. Both are impressive, and both may explain very little.
A serious research platform must therefore preserve a complete research lineage: the original hypothesis, every test definition, the data version, the parameter choices, and the failed results should all leave a record. Failed research is not garbage that needs to be cleaned up. It is the institution's error-prevention asset.
A team that keeps only successful backtests will systematically overestimate its own talent. What it actually operates is not a research database, but a museum of survivorship bias.
The ordinary problems of alternative data
Traditional data are not novel. Prices, volumes, financial statements, regulatory filings, and macro indicators are easy to obtain and have already been analyzed by countless investors. Their advantage is that definitions are relatively stable, history can be traced, and audit is easier; their disadvantage is that uniqueness is limited.
Alternative data promise an earlier or finer observation of economic activity. News, web search, social media, and other non-traditional sources may reveal changes in demand, competition, or management tone before formal financial reports.
But "alternative" is only a classification of source, not a certification of quality.
The history of alternative data is usually shorter, and coverage may be biased toward certain firms, languages, or regions. Website structures change, vendors backfill records, and the same news item may also be counted more than once. The most dangerous errors often come from time. The publication time in a research database is not necessarily the time at which investors could actually obtain the information.
Whether alternative data can produce trading value depends on three moments: when the information appears, when the institution can process it reliably, and when the market completes the pricing.
If the market absorbs the news within two hours, while the data team needs two days to finish scraping, entity matching, and cleaning, then unique data will produce only a unique historical explanation. They will not produce a tradable edge.
Machines that can read reports
The most realistic contribution of generative AI to factor research is not predicting stock prices, but compressing the cost of reading.
Annual reports, management statements, and news can be converted into structured fields. Models can extract demand outlooks, capital-expenditure plans, the competitive environment, risk wording, and strategic priorities, and then map this content into comparable labels.
This is valuable, but three things must be distinguished.
The first is information extraction: what the document said. The second is standardized classification: whether these statements belong to expansion, maintenance, or contraction, and whether risk is rising or falling. Only the third is an investment forecast: whether these labels can predict future returns, earnings revisions, volatility, or credit risk.
The first two belong to semantic processing. Only the third belongs to factor research.
That a piece of news is judged positive by a model does not mean the stock price will rise. The market may have expected better news, valuation may already reflect an optimistic scenario, and positive fundamentals may also be accompanied by a higher discount rate. Sentiment is sometimes only an echo of a price change, not the cause.
Professional research should therefore not be satisfied with an absolute sentiment score. A more meaningful object is surprise sentiment, that is, the deviation of the current tone from market expectations; one can also study changes in tone, disagreement across sources, and inconsistency between management language and financial numbers.
Generative AI can turn text into data. It cannot automatically turn data into Alpha.
A referee that drifts
Using large language models to generate factors also introduces a problem that traditional financial data encounter less often: the scorer itself will change.
Repeating the processing of the same text may produce different results. A model upgrade will change historical scores, and a small adjustment to the prompt may also move the classification boundary. Truncation of long documents, negation, legal wording, and cross-language differences may all cause systematic misjudgment. If the input document contains a malicious instruction, the model may even deviate from the assigned task.
The definition of a semantic factor is therefore not only a formula. It also includes the foundation model, the prompt, the input version, the decoding settings, and the classification system.
A change in any one of these may be equivalent to replacing the factor.
An institution should not simply append new-model scores to the end of an old historical series. It needs to preserve the original documents, the time of acquisition, the model responses, and complete version information, and to evaluate the migration between old and new scores. When necessary, it should reprocess the entire history and then repeat the out-of-sample and portfolio tests.
A software team may treat a model upgrade as maintenance. An investment team should treat it as a change in the research hypothesis.
Many versions of the same fact
Financial institutions rarely lack data. They more often lack consistent data.
Trading databases, data warehouses, data lakes, search systems, feature stores, and researchers' local files often store different versions of the same fact. A simple debt-to-equity ratio may use total debt, net debt, current liabilities, end-of-period equity, or average equity. If a research report writes only "the D/E factor," the result is almost impossible to reproduce.
A factor catalog therefore needs to record the precise formula, data fields, source, calculation frequency, reporting lag, missing-value rules, outlier treatment, industry-neutralization method, investment universe, version, owner, and approval status.
The most important function of a modern data platform is not to put everything into one enormous storage pool, but to enable another researcher to reproduce the same result using the data that were actually available at the time and the same version of the code.
Reproducibility sounds unglamorous. Irreproducible Alpha is usually even less valuable.
The cloud is not a larger hard drive
When financial institutions build a factor platform, they easily treat the architecture diagram as progress itself. Data lakes, data warehouses, stream processing, vector search, graph databases, and machine-learning services are placed on the same page; the more arrows there are, the more modern the project seems.
Technology selection should begin from access patterns, not from a product list.
Historical prices and financial data suit large-scale analysis; news and documents need text and semantic retrieval; supply-chain, management, and equity relationships suit graph queries; real-time messages need stream processing; hot features need low-latency reads; and cross-market historical replay needs elastic distributed computing.
On AWS, this may mean using Amazon S3 and an open table format to store auditable data, Amazon Redshift to carry analytical queries, Amazon EMR to process large-scale data, Amazon OpenSearch Service to support text and vector retrieval, Amazon Neptune to handle relationship networks, Amazon Timestream to store time series, Amazon SageMaker AI to train and deploy models, and AWS Lake Formation together with catalog services to unify permissions and governance.
The value of these services does not lie in their number, but in the fact that each service corresponds to a clear requirement of latency, scale, recovery, and regulation.
An architecture diagram cannot produce Alpha. At most, it can reduce the time required to produce and discover errors.
The political economy of the lakehouse
The real conflict in data architecture is not about storage format. It is between research freedom and institutional control.
Researchers want to obtain new data quickly, build experimental tables, use different engines, and share preliminary results. Risk, compliance, and audit teams need to know whether the data were authorized, which version entered the model, how the factor was calculated, who accessed the materials, and which portfolios an error would affect.
If governance is too heavy, researchers will bypass the platform. If governance is too light, research results cannot enter production.
The value of a lakehouse architecture is to use a unified catalog, business vocabulary, open table formats, data lineage, and fine-grained permissions to turn governance into research infrastructure, rather than an approval barrier after research is finished.
The best guardrail does not stop a car from moving forward. It only lowers the probability that the car will leave the road. Data governance should be the same.
Significant, yet not necessarily important
Beta, t-statistics, and R² are research tools, not a capital-allocation committee.
Beta measures a security's sensitivity to a given factor, but high sensitivity does not mean the factor can predict returns. R² indicates how much in-sample variation the model explained, yet it does not show that the model possesses independent Alpha. Many valuable cross-sectional signals can explain only a small part of short-horizon volatility; a model highly exposed to market direction may, by contrast, have a very high R².
The t-statistic is also easily granted excessive respect. Autocorrelation, cross-sectional correlation, overlapping holding periods, multiple testing, and sample selection may all inflate significance. More important, statistical significance does not mean that something is economically worth trading.
A factor must pass at least four gates.
The first is statistical validity, including confidence intervals, stability, and multiple-testing corrections. The second is predictive validity, including out-of-sample IC, ranking monotonicity, and cross-market performance. The third is economic validity: whether the return has a coherent mechanism and is independent of known risk factors. The fourth is trading validity, including turnover, spreads, market impact, securities borrowing, financing, and capacity.
Statistics can say that a result does not look much like chance. It will not pay the trader's commissions.
The trial of a leverage factor
The debt-to-equity ratio is a good example of a research trap.
The simplest backtest calculates the ratio for each firm, sorts names into high and low groups, and then compares the return of the long-short portfolio. Such a result may look clear, yet the economic meaning is not clear.
High leverage may represent financial fragility, or it may represent an efficient capital structure, the value of a tax shield, or an industry business model. The normal leverage of banks, utilities, real estate, and asset-light technology firms cannot be compared directly. Without industry adjustment, a so-called leverage factor may only be industry rotation.
The denominator will also create trouble. When shareholders' equity is near zero or negative, the ratio may become extreme or even lose meaning. The report date and the date on which investors actually obtain the data are also different; if the fiscal-period-end date is used in place of the disclosure date, the backtest will peek into the future.
Highly leveraged firms are also often jointly exposed to small size, low quality, high volatility, low liquidity, and credit risk. Without multi-factor attribution, the researcher may only have given a known risk a new name.
The real question is therefore not whether the debt-to-equity ratio was correlated with historical returns, but whether, after information becomes public, after industry and known factors are controlled, and after transaction costs are deducted, it still provides stable incremental predictive value.
That sentence is much longer than a ranking backtest, and much closer to investment reality.
Three risks in twenty factors
The number of factors is not equal to diversification.
PB, PE, PEG, and cash-flow yield may all be expressing valuation. RSI, ROC, past returns, and news sentiment may be jointly exposed to short-horizon momentum. If a team counts factors by name, it will overestimate the number of genuinely independent bets.
Generative AI will also manufacture surface richness. "CEO sentiment," "strategic tone," "risk wording," and "news sentiment" can have different names, yet all depend on the same foundation model, prompt template, and text distribution. Once the model is upgraded or the language environment changes, they may drift at the same time.
Genuine diversification requires not only low historical correlation, but also different economic sources, data dependencies, and failure paths. Portfolio construction should inspect signal correlation, position overlap, stress-period correlation, shared transaction costs, and shared model dependence.
Five factors with different names may still be a single crowded trade.
Risk has four layers
The risk of a factor strategy does not exist only in prices.
The first layer is data risk, including delay, field changes, vendor interruption, duplicated news, timestamps, and security-mapping errors. The second layer is model risk, including overfitting, look-ahead bias, parameter instability, feature leakage, and LLM output drift. The third layer is portfolio risk, including market Beta, industry concentration, liquidity, leverage, short squeezes, and tail correlation. The fourth layer is platform risk, including data-pipeline failure, permission errors, deployment of the wrong version, duplicated order submission, and monitoring failure.
Strategy status therefore cannot be determined by P&L alone. The institution must be able to judge whether a loss comes from signal failure, deteriorating execution, a data error, a model deployment, or a shock to a common risk factor.
The most dangerous thing is not a loss. It is a team that can see only the loss and cannot explain where it came from.
Factors also need funerals
A successful factor accumulates research reputation, technical investment, and management expectation. The team builds systems around it and also binds personal career achievement to it. At that point, admitting that the factor has failed is no longer only a statistical judgment. It also becomes an organizational loss.
A retirement regime should therefore be defined when the factor goes live.
When predictive power declines continuously, it should enter a watch list. When transaction costs, correlation, or crowding rise, its weight should be reduced. When data integrity cannot be confirmed, when model output shows abnormal drift, or when live execution diverges severely from the simulation, trading should be paused. If out-of-sample predictive power disappears for a long period, if after-cost return remains persistently negative, or if the original other side of the trade no longer exists, it should be retired permanently.
Closing a factor does not mean that the research failed. Maintaining the original capital allocation after the evidence has changed is the failure.
An excellent factor factory has not only a production line, but also a scrap-handling system.
Moving servers will not move Alpha in
Cloud migration is often misunderstood as putting local servers into someone else's data center. This can reduce hardware management, yet it will not automatically solve data silos, manual deployment, inconsistent research environments, unreproducible models, and uncontrolled permissions.
Genuine modernization should rebuild the research operating system.
The institution first inventories data, models, dependencies, licenses, and risk requirements, then establishes multiple accounts, unified identity, centralized logs, and permission boundaries. It next migrates low-risk data and batch processing, builds standard pipelines, a model registry, automated testing, and monitoring, and finally splits tightly coupled processes into event-driven data ingestion, quality checks, feature computation, backtesting, risk approval, and signal publication.
Whether the migration succeeded should not be measured by how many servers were moved, but by how much time, how many manual handoffs, and how many unreproducible steps are required from proposing a factor hypothesis to obtaining an auditable out-of-sample result.
A technology migration that does not change research discipline is only putting an old problem onto a new bill.
A minimum product still needs complete responsibility
A factor platform is suited to being built by the MVP method, but minimum viable does not mean minimum control.
A meaningful MVP should choose a clear hypothesis, connect one set of traditional data and one set of text data, establish traceable ingestion, compute traditional and semantic factors, complete out-of-sample tests, construct a simple portfolio, add a cost model, and then observe live performance through a shadow portfolio.
It must cover a complete vertical slice from data to risk monitoring.
A data lake without research results is not a factor platform. A sentiment model without out-of-sample validation is also not a factor platform. A backtest without versioning, monitoring, and a shutdown mechanism is even less able to enter production.
In a financial system, "fast" should not mean skipping controls. It should mean automating the necessary controls.
Guardrails for research freedom
A factor platform cannot be built by an engineering team alone. Investment managers, quantitative researchers, data engineers, model risk, market risk, compliance, legal, security, audit, and FinOps will all take part.
Organizations easily move toward two extremes. If central governance requires every experiment to pass through a long approval, researchers will bypass the platform. If research teams are entirely free to choose data, definitions, and deployment methods, the production system will be unauditable and costs will also run out of control.
A more reasonable model is for the central team to establish the minimum necessary standards, and for researchers to experiment on a self-service basis inside the guardrails. A model that affects real capital must have a version, lineage, an owner, monitoring, and an independent shutdown right. The higher the risk, the stricter the approval; the earlier the experiment, the greater the freedom.
The cultural principle is simple: researchers may make mistakes quickly, but the institution must be able to reproduce, explain, constrain, and close those mistakes.
Producing errors at ten times the speed
A modern factor platform can easily look advanced. Streaming ingestion, open table formats, vector databases, graph databases, large language models, multi-agent workflows, and real-time dashboards can make an architecture demonstration highly persuasive.
What these tools raise is research throughput, not the truth rate of research.
If an institution raises the speed of factor search by a factor of ten, yet does not improve multiple testing, out-of-sample validation, cost estimation, and the retirement regime, it has only obtained the ability to manufacture false Alpha at ten times the speed.
The value of technology should be measured by more modest metrics: whether experiments can be reproduced, how quickly data errors are discovered, what the out-of-sample pass rate is, how long research takes to enter a shadow portfolio, whether production models can be traced, how quickly a failed factor is closed, and whether capital can be reallocated in time.
The highest aim of a factor platform is not to make backtests more handsome, but to lower the probability that the institution will allocate capital to a wrong model.
The production function of Alpha
Modern factor competition is shifting from "who has more data" to "who can complete high-quality learning faster."
Traditional data provide a stable benchmark, alternative data contest the information-time gap, generative AI compresses unstructured information, and the cloud platform provides elastic compute and unified governance. But these tools can produce sustainable value only when they are combined with economic hypotheses, statistical discipline, transaction costs, portfolio risk, and model governance.
An institution should not ask how many factors it has discovered. It should ask how many factors have an explainable source of return, pass out-of-sample tests, still have positive incremental value after costs, and can be identified and closed in time when they fail.
A genuine factor factory does not define success by output. It defines success by how quickly errors are discovered, how robustly an edge is deployed, and how promptly failed capital is withdrawn.
The cloud can accelerate computation, and AI can accelerate reading. Neither can replace an older investment capability: knowing when an exciting number is still only a number.