引言
回归是监督学习里最直观、也最实用的一类问题:给定若干特征(面积、房龄、位置),预测一个连续数值(房价)。线性回归假设「输出是输入的线性组合」,这个简单假设能解决大量真实问题,也是理解更复杂模型(逻辑回归、神经网络)的起点。
本文不满足于「调一行 fit()」:先从最小二乘的数学直觉讲起,让你知道模型在优化什么;再用真实结构化的房价数据跑通一元回归、多元回归、多项式回归,观察过拟合长什么样;接着引入正则化(岭回归、Lasso)与交叉验证,最后用评估指标(MSE、MAE、R²)判断模型好坏。
前置:[[ml]] 专题的环境搭建与 pandas 基础(https://plumephp.com/ml-python-environment-setup/)。本专题聚焦「动手跑通 + 直觉理解」,更严谨的推导与更广的回归变体见 [[ai-ml]] 专题。
目录
- 1. 回归问题与最小二乘原理
- 2. 数据准备:真实房价数据集
- 3. 一元线性回归:第一个模型
- 4. 多元线性回归与特征组合
- 5. 数据标准化:让特征站在同一起跑线
- 6. 多项式回归:当直线不够用时
- 7. 过拟合与正则化:岭回归与 Lasso
- 8. 交叉验证与模型评估
- 9. 总结:回归问题的完整套路
- 延伸阅读
1. 回归问题与最小二乘原理
1.1 回归建模直觉
给每个样本一个预测 y_hat = w1*x1 + w2*x2 + ... + b,我们的目标是找到一组 w(权重)和 b(偏置),让预测尽量接近真实值。
1.2 最小二乘:最小化误差平方和
损失函数(MSE)是所有样本预测误差平方的平均:
L(w, b) = (1/N) * Σ (y_i - y_hat_i)²
「训练」就是找到使 L 最小的 w 和 b。最小二乘给出闭式解,但更常用的是梯度下降:从初始点出发,沿损失函数下降最快的方向(负梯度)一步步走。
1.3 梯度下降一句话
w ← w - 学习率 × 损失对 w 的偏导
| 概念 | 直觉 |
|---|---|
| 损失 | 模型有多「错」 |
| 梯度 | 往哪个方向调,错得更快减少 |
| 学习率 | 每步走多大 |
2. 数据准备:真实房价数据集
2.1 用 sklearn 内置的加州房价数据集
from sklearn.datasets import fetch_california_housing
import pandas as pd
housing = fetch_california_housing(as_frame=True)
df = housing.frame
df['MEDV'] = df['MedHouseVal'] # 目标:房价中位数
print(df.head())
print(df.shape) # (20640, 9)
2.2 快速探索
print(df.describe()) # 各列分布
print(df.isnull().sum()) # 缺失值
# 特征与目标的相关性
corr = df.drop(columns=['MedHouseVal']).corrwith(df['MedHouseVal'])
print(corr.sort_values(ascending=False))
2.3 特征/目标切分
features = ['MedInc', 'HouseAge', 'AveRooms', 'AveBedrms',
'Population', 'AveOccup', 'Latitude', 'Longitude']
X = df[features]
y = df['MEDV']
3. 一元线性回归:第一个模型
3.1 只用「收入中位数」预测房价
from sklearn.model_selection import train_test_split
from sklearn.linear_model import LinearRegression
X1 = X[['MedInc']] # 单特征
X_train, X_test, y_train, y_test = train_test_split(
X1, y, test_size=0.2, random_state=42)
model = LinearRegression()
model.fit(X_train, y_train)
print("系数 w:", model.coef_[0]) # 收入每+1,房价约+0.42
print("截距 b:", model.intercept_)
3.2 画拟合直线与残差
import matplotlib.pyplot as plt
y_pred = model.predict(X_test)
plt.scatter(X_test, y_test, alpha=0.4, label='真实')
plt.scatter(X_test, y_pred, alpha=0.4, label='预测')
plt.xlabel('MedInc (收入)')
plt.ylabel('房价中位数')
plt.legend()
plt.show()
4. 多元线性回归与特征组合
4.1 用全部特征
X_train, X_test, y_train, y_test = train_test_split(
X, y, test_size=0.2, random_state=42)
model = LinearRegression()
model.fit(X_train, y_train)
# 打印每个特征的重要性(粗略)
for name, coef in zip(features, model.coef_):
print(f"{name:12s} {coef:+.4f}")
4.2 观察:多个特征是否一定更好
多元回归通常优于一元,但不代表特征越多越好——冗余与噪声特征反而增加方差。下一节我们用标准化与正则化来治理。
5. 数据标准化:让特征站在同一起跑线
5.1 为什么需要标准化
MedInc 量级在个位,Population 在千位。量级大的特征会在距离计算与正则化中「霸占」权重。标准化把每个特征变成均值 0、标准差 1。
from sklearn.preprocessing import StandardScaler
scaler = StandardScaler()
X_scaled = scaler.fit_transform(X) # 拟合后转换训练数据
# 标准化后各特征均值≈0,标准差≈1
print(X_scaled.mean(axis=0).round(3))
print(X_scaled.std(axis=0).round(3))
5.2 必须用训练集拟合、测试集转换
scaler = StandardScaler()
X_train_s = scaler.fit_transform(X_train) # 只用训练集学参数
X_test_s = scaler.transform(X_test) # 测试集只转换,不重新拟合
这是防止数据泄漏的基本纪律:测试集的信息不能进入训练过程。
6. 多项式回归:当直线不够用时
6.1 生成非线性数据
import numpy as np
np.random.seed(42)
x = np.linspace(0, 10, 200)
y_true = x**2 - 3*x + 5
y = y_true + np.random.normal(0, 15, 200)
6.2 用 PolynomialFeatures 提升到二次
from sklearn.preprocessing import PolynomialFeatures
from sklearn.pipeline import make_pipeline
pipe = make_pipeline(
PolynomialFeatures(degree=2), # 构造 x, x²
LinearRegression(),
)
pipe.fit(x.reshape(-1, 1), y)
score = pipe.score(x.reshape(-1, 1), y)
print("二次拟合 R²:", round(score, 3))
6.3 degree 越高越容易过拟合
for degree in [1, 2, 10, 20]:
pipe = make_pipeline(
PolynomialFeatures(degree=degree),
LinearRegression(),
)
pipe.fit(x.reshape(-1, 1), y)
# 训练集 R² 会接近 1,但测试集可能很差(过拟合)
print(f"degree={degree:2d} 训练R²={pipe.score(x.reshape(-1,1), y):.4f}")
degree 太大,模型把噪声也「背」下来——这就是过拟合:训练很好、泛化很差。
7. 过拟合与正则化:岭回归与 Lasso
7.1 正则化思想
在损失里加一项惩罚,约束权重不要太大:
岭回归: L + α * Σ w² (惩罚平方和,权重整体缩小)
Lasso: L + α * Σ |w| (惩罚绝对值,可把部分权重压到 0)
α 越大,惩罚越强,模型越简单。
7.2 岭回归实战
from sklearn.linear_model import Ridge, Lasso
ridge = Ridge(alpha=1.0)
ridge.fit(X_train_s, y_train)
print("岭回归 R²:", ridge.score(X_test_s, y_test))
7.3 Lasso 特征选择
lasso = Lasso(alpha=0.1, max_iter=10000)
lasso.fit(X_train_s, y_train)
for name, coef in zip(features, lasso.coef_):
print(f"{name:12s} {coef:+.4f}") # 不重要的特征权重被压到接近 0
7.4 正则化系数如何选:交叉验证
from sklearn.linear_model import RidgeCV
ridge_cv = RidgeCV(alphas=[0.1, 1, 10, 100])
ridge_cv.fit(X_train_s, y_train)
print("最优 alpha:", ridge_cv.alpha_)
8. 交叉验证与模型评估
8.1 为什么不能只看一次划分
一次划分运气成分大(测试集恰好简单/难)。交叉验证把数据分 K 折,轮流当测试,取平均。
from sklearn.model_selection import cross_val_score
model = Ridge(alpha=1.0)
scores = cross_val_score(model, X_train_s, y_train,
cv=5, scoring='r2')
print("5折 R²:", scores.round(3))
print("平均 R²:", scores.mean().round(3))
8.2 回归评估指标
| 指标 | 直觉 | 越小越好 |
|---|---|---|
| MSE | 误差平方平均,放大离群点 | 是 |
| MAE | 误差绝对值平均,更鲁棒 | 是 |
| R² | 模型解释了多大比例方差 | 越大越好(≤1) |
from sklearn.metrics import mean_squared_error, mean_absolute_error, r2_score
y_pred = ridge.predict(X_test_s)
print("MSE:", mean_squared_error(y_test, y_pred).round(3))
print("MAE:", mean_absolute_error(y_test, y_pred).round(3))
print("R² :", r2_score(y_test, y_pred).round(3))
8.3 残差分析:模型假设是否成立
resid = y_test - y_pred
plt.scatter(y_pred, resid, alpha=0.4)
plt.axhline(0, color='red', linestyle='--')
plt.xlabel('预测值')
plt.ylabel('残差')
plt.show()
# 理想:残差随机分布在 0 两侧,无明显模式
9. 总结:回归问题的完整套路
9.1 五步流水线
# 完整套路(模板)
# 1. 切分 train_test_split
# 2. 标准化 StandardScaler.fit_transform(训练) / transform(测试)
# 3. 训练 模型.fit(X_train_s, y_train)
# 4. 调参/评估 交叉验证选 α,MSE/MAE/R²
# 5. 诊断 残差图、过拟合检查
9.2 关键决策点
| 问题 | 选择 |
|---|---|
| 线性关系? | 线性回归 |
| 非线性但已知形式? | 多项式回归(degree 适中) |
| 特征多/过拟合? | 岭回归 / Lasso |
| α 怎么定? | 交叉验证选最优 |
| 特征量级差异大? | 必须先标准化 |
9.3 一句话心法
回归的本质是「用数据学习一条线/一个面」,正则化与交叉验证是防止「学过头」。 先跑通线性,再逐步加复杂度,并始终用测试集评价真实能力。
延伸阅读
- https://plumephp.com/ml-python-environment-setup/ — 环境与 pandas 基础
- https://plumephp.com/ml-supervised-classification/ — 下一个任务:分类问题
- [[ai-ml]] 专题的模型评估与特征工程深度文章
- scikit-learn 线性模型文档
继续阅读
探索更多技术文章
浏览归档,发现更多关于系统设计、工具链和工程实践的内容。