提供丰富的素材资源、软件工具、源码模板、技术文章和编程教程，专注于网站搭建、AI应用、开源项目分享和工具推荐。帮助开发者轻松获取所需资源，快速提升技术水平。

搜索运维相关内容

热词：

youtube会员

YouTuBe

disney会员

Disney

Netflix奈飞账号

Netflix

iCloud+

iCloud+

hbo+max

HBOMax

GPT+API

GPTPro

Spotify会员

Spotify

合租&账号

艾维正版

莱卡云服务器

bandwagonhost云主机

雨云服务器

Python 中 set 是什么？为何要是用它？

2025-01-24 17:46

80

标签导航：

python 中 set 是什么？为何要是用它？

Python 提供多种内置数据结构用于组织数据，包括列表、字典、元组和集合。

根据 Python 3 文档，集合是无序的、不包含重复元素的集合。其主要用途包括成员测试和去除重复项。集合还支持集合运算，如并集、交集、差集和对称差集。

本文将通过示例阐述以上定义中的每个特性，并讲解集合的创建方法。

**集合的初始化**

创建集合有两种方法：一是将元素列表传递给内置函数 `set()`，二是使用花括号 `{}`。

使用 set() 函数初始化集合：

>>> s1 = set([1, 2, 3])
>>> s1
{1, 2, 3}
>>> type(s1)
<class 'set'>

使用 {} 初始化集合：

>>> s2 = {3, 4, 5}
>>> s2
{3, 4, 5}
>>> type(s2)
<class 'set'>

两种方法均有效。但如果需要创建空集合呢？

>>> s = {}
>>> type(s)
<class 'dict'>

使用空花括号将创建一个字典，而非集合。

为简便起见，本文示例使用整数集合，但集合可包含所有 Python 支持的可哈希hashable[1]数据类型，例如整数、字符串和元组，但不包括列表或字典等可变类型。

>>> s = {1, 'coffee', [4, 'python']}
Traceback (most recent call last):
  File "<stdin>", line 1, in <module>
TypeError: unhashable type: 'list'

了解了集合的创建方法及其可包含的元素类型后，我们来探讨一下为什么应该将集合纳入你的编程工具箱。

**为什么需要使用集合**

编程中，解决问题的方法不止一种。有些方法效率低下，而另一些则清晰、简洁、易于维护，或者更“Pythonicpythonic[2]”。

根据 Python Hitchhiker's Guide[3]：

当经验丰富的 Python 开发者（PythonistaPythonista）遇到不够“Pythonicpythonic”的代码时，通常认为这些代码不符合通用指南，表达意图的方式不够理想（可读性差）。

让我们看看集合如何提升代码可读性并提高程序执行效率。

**集合元素的无序性**

集合元素无法通过索引访问：

>>> s = {1, 2, 3}
>>> s[0]
Traceback (most recent call last):
  File "<stdin>", line 1, in <module>
TypeError: 'set' object does not support indexing

也无法使用切片修改：

>>> s[0:2]
Traceback (most recent call last):
  File "<stdin>", line 1, in <module>
TypeError: 'set' object is not subscriptable

但如果需要去除重复项或进行集合运算（如并集），则应该使用集合。

迭代时，集合的效率优于列表。原因涉及集合的内部实现细节，感兴趣的读者可以参考以下链接：

时间复杂度[4]
set() 的实现[5]
Python 集合 vs 列表[6]
列表中使用集合的优缺点[7]

**集合不包含重复项**

过去，我经常使用 `for` 循环和 `if` 语句来检查并去除列表中的重复元素，代码如下：

>>> my_list = [1, 2, 3, 2, 3, 4]
>>> no_duplicate_list = []
>>> for item in my_list:
...     if item not in no_duplicate_list:
...             no_duplicate_list.append(item)
...
>>> no_duplicate_list
[1, 2, 3, 4]

或者使用列表推导式：

>>> my_list = [1, 2, 3, 2, 3, 4]
>>> no_duplicate_list = []
>>> [no_duplicate_list.append(item) for item in my_list if item not in no_duplicate_list]
[None, None, None, None]
>>> no_duplicate_list
[1, 2, 3, 4]

现在，我们可以使用集合：

>>> my_list = [1, 2, 3, 2, 3, 4]
>>> no_duplicate_list = list(set(my_list))
>>> no_duplicate_list
[1, 2, 3, 4]

使用 timeit 模块比较列表和集合去除重复项的效率：

使用集合比列表推导式更简洁，效率更高。

注意：集合是无序的，转换为列表后元素顺序可能发生变化。

Python 之禅[8]：

优美胜于丑陋Beautiful is better than ugly.

明了胜于晦涩Explicit is better than implicit.

简洁胜于复杂Simple is better than complex.

扁平胜于嵌套Flat is better than nested.

集合正是如此。

**成员测试**

使用 `if` 语句检查元素是否存在于列表中，称为成员测试：

my_list = [1, 2, 3]
>>> if 2 in my_list:
...     print('Yes, this is a membership test!')
...
Yes, this is a membership test!

集合的成员测试效率更高：

>>> from timeit import timeit
>>> def in_test(iterable):
...     for i in range(1000):
...             if i in iterable:
...                     pass
...
>>> timeit('in_test(iterable)',
... setup="from __main__ import in_test; iterable = list(range(1000))",
... number=1000)
12.459663048726043

>>> from timeit import timeit
>>> def in_test(iterable):
...     for i in range(1000):
...             if i in iterable:
...                     pass
...
>>> timeit('in_test(iterable)',
... setup="from __main__ import in_test; iterable = set(range(1000))",
... number=1000)
.12354438152988223

注意：以上测试来自 StackOverflow[9]。

对于大型列表，将其转换为集合可以提高成员测试效率。

**如何使用集合**

了解了集合及其用途后，我们来快速浏览一下集合的修改和操作方法。

**添加元素**

根据需要添加的元素数量，选择 `add()` 或 `update()` 方法。

add() 用于添加单个元素：

>>> s = {1, 2, 3}
>>> s.add(4)
>>> s
{1, 2, 3, 4}

update() 用于添加多个元素：

>>> s = {1, 2, 3}
>>> s.update([2, 3, 4, 5, 6])
>>> s
{1, 2, 3, 4, 5, 6}

集合会自动去除重复元素。

**移除元素**

如果需要在删除不存在的元素时引发异常，使用 `remove()`；否则，使用 `discard()`：

>>> s = {1, 2, 3}
>>> s.remove(3)
>>> s
{1, 2}
>>> s.remove(3)
Traceback (most recent call last):
  File "<stdin>", line 1, in <module>
KeyError: 3

discard() 不会引发异常：

>>> s = {1, 2, 3}
>>> s.discard(3)
>>> s
{1, 2}
>>> s.discard(3)
>>> # 无任何操作

pop() 用于随机移除一个元素：

>>> s = {1, 2, 3, 4, 5}
>>> s.pop()  # 移除一个任意元素
1
>>> s
{2, 3, 4, 5}

clear() 用于清空集合：

>>> s = {1, 2, 3, 4, 5}
>>> s.clear()  # 清空集合
>>> s
set()

**并集 `union()`**

`union()` 或 `|` 操作符创建包含所有元素的新集合：

>>> s1 = {1, 2, 3}
>>> s2 = {3, 4, 5}
>>> s1.union(s2)  # 或 's1 | s2'
{1, 2, 3, 4, 5}

**交集 `intersection()`**

`intersection()` 或 `&` 操作符返回集合的公共元素：

>>> s1 = {1, 2, 3}
>>> s2 = {2, 3, 4}
>>> s3 = {3, 4, 5}
>>> s1.intersection(s2, s3)  # 或 's1 & s2 & s3'
{3}

**差集 `difference()`**

`difference()` 或 `-` 操作符创建新集合，包含在 `s1` 中但在 `s2` 中不存在的元素：

>>> s1 = {1, 2, 3}
>>> s2 = {2, 3, 4}
>>> s1.difference(s2)  # 或 's1 - s2'
{1}

**对称差集 `symmetric_difference()`**

`symmetric_difference()` 或 `^` 操作符返回集合中不同的元素：

>>> s1 = {1, 2, 3}
>>> s2 = {2, 3, 4}
>>> s1.symmetric_difference(s2)  # 或 's1 ^ s2'
{1, 4}

**结论**

本文介绍了 Python 集合的概念、操作方法和应用场景。熟练掌握集合的使用可以编写更清晰、更高效的代码。

如有任何疑问，请随时提出。此外，Python 速查表[10]中也包含了集合的相关内容[11]，方便查阅。

相关文章推荐

Linux JS日志如何解读

Win11 24H2部分用户反馈开机卡死/崩溃:疑和沙盒功能有关

Linux DHCP安全设置：如何保护DHCP服务

如何用Compton提升游戏体验

Linux系统垃圾清理：这些文件夹别忽视

如何用Compton优化Linux图形界面

Linux DHCP客户端配置：如何获取IP地址

Linux strings命令如何结合其他工具使用

XRender在Linux系统中的配置方法

如何快速解决Linux backlog

SecureCRT如何配置SSH密钥

SecureCRT如何实现会话共享