Every piece in this price-band series so far has used a snapshot: where a stock's price sits today. This piece asks the question that snapshot can't answer — how much do stocks actually move between price bands over time, and does an algorithm that doesn't know about price bands at all, given only price, volatility, return and beta, group stocks any differently than the manual bands this blog has been drawing? Both analyses here are restricted to confirmed equity tickers only, cross-checked against this series' own Screener.in price-band lists — an earlier version of this piece drew from the full local price warehouse without that filter and was skewed by corporate bonds trading near their ₹1–10 lakh face value; see the correction note at the end.
Do Cheap Stocks Stay Cheap? Price-Band Migration and What K-Means Clustering Actually Finds
1. How much migration actually happens between price bands
Restricted to the 5,390 equity tickers used throughout this series (not the raw local warehouse, which also carries thousands of corporate bonds), this article tracked each stock's price band at the start of the window (January 2023) against its band at the end (August 2026). Local price history wasn't available for every ticker over the full 3.5-year span, so the table below shows both the tracked count and the coverage rate against each band's full Screener.in universe:
| Starting price band, ₹ | Stocks tracked (of Screener universe) | Coverage, % | Moved to a different band, % | Moved up, % | Moved down, % |
|---|---|---|---|---|---|
| 0-50 | 555 of 2,086 | 26.6 | 18.2 | 18.2 | 0.0 |
| 50-100 | 332 of 743 | 44.7 | 65.7 | 46.4 | 19.3 |
| 100-150 | 233 of 469 | 49.7 | 73.0 | 47.6 | 25.3 |
| 150-250 | 306 of 552 | 55.4 | 63.7 | 44.4 | 19.3 |
| 250-500 | 437 of 568 | 76.9 | 63.2 | 44.6 | 18.5 |
| 500-750 | 189 of 290 | 65.2 | 69.8 | 47.1 | 22.8 |
| 750-1,000 | 105 of 154 | 68.2 | 83.8 | 57.1 | 26.7 |
| 1,000-3,000 | 221 of 365 | 60.5 | 44.8 | 30.8 | 14.0 |
| 3,000-5,000 | 34 of 73 | 46.6 | 70.6 | 58.8 | 11.8 |
| 5,000-10,000 | 20 of 52 | 38.5 | 45.0 | 30.0 | 15.0 |
| 10,000+ | 12 of 38 | 31.6 | 16.7 | 0.0 | 16.7 |
With bonds removed, the pattern is sharper than the uncorrected version suggested. The ₹0–50 band is still the most immobile — 18.2% moved, and every move was upward, since there's nowhere cheaper to go — but every other band shows real, often majority-level churn: 63–84% of stocks starting anywhere between ₹50 and ₹1,000 ended the window in a different band. Coverage is real limitation worth stating plainly rather than glossing over: it ranges from 26.6% (₹0–50, the band with the most thinly-traded, sparsely-covered names) up to 76.9% (₹250–500), and only 12 of the 38 stocks in the above-₹10,000 band had a complete local price history to track — a small enough base that the 16.7%-moved figure for that band should be read as illustrative, not a precise population estimate.
2. What K-means clustering finds when it isn't told about price bands
Same correction applied here: K-means clustering (four clusters, standardised features) was re-run using only the 1,761 confirmed equity tickers with sufficient price history — a stricter bar than the migration table's 2,444, since clustering needs a continuous price series across the full 3.5-year window to compute annualised return, volatility, and beta, while migration only needs a start and end price — on four inputs — price (log-scaled), annualised return, annualised volatility, and beta — with the algorithm given no knowledge of this series' own ₹250/₹1,000/₹10,000 boundaries.
Each point is one equity stock: x-axis is price (log scale), y-axis is annualised volatility (%). Colour is the K-means cluster assignment.
| Cluster | Character | Stocks | Median price, ₹ | Median return, % p.a. | Median volatility, % p.a. | Median beta |
|---|---|---|---|---|---|---|
| Cluster 0 | Cheap & distressed (lowest median return) | 368 | 22.36 | -14.76 | 52.09 | 1.31 |
| Cluster 3 | Ordinary mid-price | 527 | 240.6 | 10.85 | 45.0 | 1.41 |
| Cluster 1 | Stable / defensive (lowest volatility & beta) | 655 | 603.25 | 9.0 | 33.3 | 0.91 |
| Cluster 2 | High-growth / momentum (highest median return) | 211 | 536.1 | 61.76 | 54.56 | 1.39 |
The corrected clustering still agrees with this series' core finding. The cheap-and-distressed cluster (median price ₹22) lines up with the ₹0–50 band's negative-return, high-beta profile found throughout this series. And the same key nuance survives the correction: the defensive cluster and the momentum cluster sit at broadly overlapping mid-range prices (medians roughly ₹540–₹600) yet have opposite risk-and-return profiles — a stock's price alone, even to an algorithm with no band boundaries to work from, does not reliably separate calm, low-beta names from volatile, high-return ones once you're above the cheapest cluster.
3. Testing K-means against two other unsupervised methods
K-means has a specific, well-documented limitation relevant here: it forces every stock into one of a pre-chosen number of roughly spherical clusters, whether or not that structure actually exists in the data, and it has no way to say "this stock doesn't really belong anywhere." Two other unsupervised methods handle that differently — DBSCAN groups points by density and explicitly flags points that don't belong to any dense region as noise, while hierarchical clustering builds a full nested tree with no cluster count decided in advance. Both were run on the same 1,761-stock equity dataset and the same four standardised features (log price, return, volatility, beta) as the K-means analysis above.
DBSCAN: one dense core, and 293 genuine outliers
Unlike K-means' four roughly-equal clusters, DBSCAN found the Indian equity market's price/risk data forms one large, dense core (1,468 stocks, median price ₹290, median return 7.4%, median beta 1.18) with 293 stocks (16.6%) flagged as noise — not close enough to any dense region to belong to a cluster at all. That's a materially different kind of finding than K-means can produce, because K-means never gets to say "this one doesn't fit anywhere" — it always assigns every point to its nearest centre regardless of fit.
Blue = the dense core cluster; red = DBSCAN noise points (genuine outliers, not assigned to any cluster). Same axes as the K-means chart above.
What makes the outlier group interesting is that it isn't a price band at all — it spans the entire range, from a ₹0.45 penny stock to MRF itself at ₹1,32,700. What the 293 outliers share isn't price; it's behaviour: a median annualised return of 28.7% against the core's 7.4%, and volatility of 56.6% against the core's 41.2%. DBSCAN is picking out stocks whose price/return/volatility/beta combination is genuinely unusual for the market as a whole, regardless of what they cost per share — a distinction K-means' forced four-way split has no mechanism to draw.
Hierarchical clustering: a second algorithm, a similar story
Agglomerative (Ward-linkage) hierarchical clustering, cut to the same four groups as the K-means analysis for comparability, arrives at a broadly similar picture through a completely different method — no centroids, no distance-to-nearest-centre assignment, just repeated merging of the two closest points or clusters until everything is one tree:
Truncated dendrogram (last 12 merges shown; leaf labels are the number of stocks folded into each branch at that point). The height of each join is the Ward-linkage distance between the two branches it merges — a taller join means a more dissimilar pair being forced together.
| Cluster | Character | Stocks | Median price, ₹ | Median return, % p.a. | Median volatility, % p.a. | Median beta |
|---|---|---|---|---|---|---|
| Cluster 3 | Cheap & distressed | 219 | 23.05 | -15.17 | 49.4 | 0.98 |
| Cluster 4 | Cheap-to-mid, flat | 453 | 91.95 | -0.47 | 48.02 | 1.51 |
| Cluster 2 | High-growth / momentum | 321 | 461.6 | 44.82 | 52.11 | 1.42 |
| Cluster 1 | Stable, largest group | 768 | 587.53 | 10.41 | 34.05 | 1.01 |
Two independent algorithms — K-means and hierarchical clustering — land on the same broad shape: a cheap, distressed group; a larger, calmer group; and a smaller high-return, high-volatility group, with a fourth group of ordinary cheap-to-mid-price stocks. That agreement across genuinely different methods is a stronger signal than either method alone: the underlying structure in the data (cheap-and-distressed vs. calm vs. momentum) doesn't depend on which clustering algorithm happens to be used to find it.
4. What the momentum literature says about the "high-growth" cluster
Both algorithms independently carve out a cluster that looks like textbook momentum: mid-priced, high beta, and by far the highest median return of any group (K-means: 211 stocks, median return 61.8% p.a.; hierarchical: 321 stocks, median return 44.8% p.a., beta 1.42). That combination — a cluster of stocks with unusually strong recent returns and above-average risk — matches decades of academic work on the "momentum" anomaly, so it's worth checking what that literature actually says rather than just applying the label.
The foundational result is Jegadeesh and Titman's 1993 paper, which found that US stocks with the strongest returns over a 3–12 month formation period kept outperforming the weakest for a further 3–12 months, and that the effect has since been replicated across international markets and asset classes over the following three decades. It was formalised as a standard risk factor in the Carhart four-factor model, and on a long-short basis (long the strongest past winners, short the weakest past losers) the momentum premium has run roughly 6–8% annually gross — a fraction of this cluster's 44.8–61.8% median return, because that figure is the cluster's own absolute return, not a winners-minus-losers spread, and because a single 3.5-year Indian window is not the same thing as a multi-decade global average.
India-specific research finds the same basic pattern domestically: a sectoral study of NSE momentum found the effect present across industries, and a study of "physical momentum" portfolios built from NSE 500 stocks found abnormal returns held up across daily, weekly, monthly and yearly holding periods. A separate portfolio-based study of Indian momentum returns also identified price-to-earnings, price-to-book and net FII inflows as factors associated with the effect — a reminder that what looks like a pure price/return/volatility cluster here may be standing in for fundamentals this piece never measured.
Two caveats the literature is explicit about, and that apply directly to this cluster's elevated 52–55% volatility and 1.4-ish beta: first, momentum's academic status is contested — Fama has argued it looks like a short-term, transient effect rather than compensation for a stable risk factor, unlike size or value; second, momentum strategies carry real tail risk, with documented crashes (most famously in 2009) when past losers stage a sharp, sudden recovery and erase a large share of the accumulated premium in a matter of weeks. Nothing in this piece's clustering tests whether these particular 211–321 stocks keep outperforming going forward — the cluster is a description of the last 3.5 years, not a forecast, and the literature above is explicit that momentum's own history includes some of its worst drawdowns immediately after periods that looked exactly like this one.
5. Reconciling the findings
Read together, the migration data and the clustering point at the same underlying story from different angles. The ₹0–50 band is the most immobile and the algorithm's own cheapest cluster is its most behaviourally distinct one — consistently distressed, high-beta, negative-return. That's a real, structural floor. Above that floor, both migration (a majority of stocks in every band from ₹50 to ₹1,000 changed bands over 3.5 years) and clustering (near-identical median prices producing opposite risk profiles) agree that price stops being a reliable organising variable well before the round-number boundaries (₹1,000, ₹10,000) this series has used to structure the analysis.
Related on this blog
See also: ₹0–50 Is Its Own Category: Diving Into India's Actual Microcap Basement · Below ₹250 Is Where the Real Risk Lives · A Slice of Apple Costs About ₹1,700. A Slice of MRF Still Isn't Legal.
Sources
- Price-band migration and K-means clustering, restricted to the 5,390-stock confirmed equity universe used throughout this series (Screener.in company data, 11 Aug 2026) — both computed from a local historical NSE OHLCV warehouse (adjusted close, year-partitioned parquet, 2 Jan 2023–7 Aug 2026)
- Nifty 50 benchmark series and risk-free proxy sourced via Yahoo Finance, cross-checked against Screener.in's Nifty 50 index page
- "DBSCAN Vs K-Means," IBKR Quant / QuantInsti — the K-means limitations (forced spherical clusters, no outlier detection, a pre-chosen cluster count) that motivated adding DBSCAN and hierarchical clustering to this piece
- Jegadeesh & Titman, "Profitability of Momentum Strategies: An Evaluation of Alternative Explanations," NBER Working Paper 7159 and the 30-years-later replication survey — the foundational momentum literature used to contextualise Cluster 2
- "Momentum Effect in Indian Stock Market: A Sectoral Study," Garg & Varshney (2015) and "Physical Momentum in the Indian Stock Market" — India-specific momentum evidence
- Ehsani & Linnainmaa, "Factor Momentum and the Momentum Factor" — on momentum's contested status as a stable risk factor and its documented crash risk
This analysis is based on historical price data cited above. It is provided for informational and research purposes only and does not constitute investment advice. Local price-history coverage varied by band (26.6–76.9% of each band's full equity universe); bands with low coverage or a small tracked count (notably above-₹10,000, at 12 stocks) should be read as illustrative rather than precise population estimates. K-means clustering is an unsupervised statistical method with no causal interpretation — it groups stocks by similarity in the four measured features over this specific window, not by any underlying economic mechanism, and cluster assignments would likely shift in a different time period.
About this article: Researched, written and edited by Umashankar Triplicane Dwarakanathan, with AI research assistance; every figure is meant to trace to the primary source cited. See the Editorial Policy for how sourcing, AI use and corrections work.